CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Fredithefish /Nemotron-CC-HQ-20B Nemotron-CC-HQ-20B This Dataset consists of approximately 20B tokens of Nemotron-CC-HQ, consisting of randomly sampled slices from crawls in the range CC-MAIN-2013-20-part-00012 to CC-MAIN-2019-04-part-00007. For more information about Nemotron-CC check the Paper by Nvidia Disclaimer: Derived from Nemotron-CC (Common Crawl). No ownership of underlying content is claimed. Data may be subject to third-party rights. Use at your own risk and in compliance with… See the full description on the dataset page: https://huggingface.co/datasets/Fredithefish/Nemotron-CC-HQ-20B.texttext-generation10M<n<100M1 likes2.1k downloads6mo agoHugging Face02ericrcwu /week1-general-20b-dolma2-v1 Week-One General 20B Dolma2 This is a deterministic, pretokenized 20-billion-token baseline corpus for controlled language-model architecture and training experiments. It contains nested 100M, 1B, 5B, and 20B views; each larger view is an exact ordered extension of the previous view. It also includes a dataset-only 370m-1.25xc view: the first 9,281,564,672 packed tokens of the verified 20B order, sized for the 1.25xC target of the OLMo-ladder 370M parameter count. The repository… See the full description on the dataset page: https://huggingface.co/datasets/ericrcwu/week1-general-20b-dolma2-v1.texttext-generationn<1K0 likes317 downloads2mo agoHugging Face03AlgoDriveAI /TinyMathStories_gpt-oss-20b TinyMathStories A TinyStories-style corpus extended with math and lightweight reasoning. This dataset keeps the child-level vocabulary and short narrative style of TinyStories (Microsoft Research, Eldan & Li, 2023) and mixes in basic numeracy (counting, addition/subtraction, simple equations, fractions, measurement) and short justifications—so tiny models can practice coherent English and early math/logic. Research, generation, and curation by AlgoDriveAI.Inspired by and… See the full description on the dataset page: https://huggingface.co/datasets/AlgoDriveAI/TinyMathStories_gpt-oss-20b.texttext-generation100K<n<1M0 likes106 downloads9mo agoHugging Face04Jackrong /GPT-OSS-20B-Distilled-Reasoning-Mini Dataset Card for Dataset Name GPT-OSS-20B Distilled Reasoning Dataset Mini (Multi-stage Evaluative Refinement Method for Reasoning Generation) Dataset Details and Description This is a high-quality instruction fine-tuning dataset constructed through knowledge distillation, featuring detailed Chain-of-Thought (CoT) reasoning processes. The dataset is designed to enhance the capabilities of smaller language models in complex reasoning, logical analysis, and instruction… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/GPT-OSS-20B-Distilled-Reasoning-Mini.tabulartext-classification1K<n<10K21 likes54 downloads1y agoHugging Face05ahmetggg /gptoss20b-bilingual-curriculum-sft gpt-oss-20b Bilingual Curriculum SFT Synthetic bilingual supervised fine-tuning data generated with gpt-oss-20b (MoE, ~3.6B active params, native MXFP4, adaptive reasoning effort by difficulty). Domains: mathematics, physics, chemistry, biology, computer science, general science, general knowledge, conversation. Languages: Turkish and English. Difficulty levels: 1-8. The dataset is synthetic and should be independently evaluated before production use. texttext-generationn<1K0 likes44 downloads1mo agoHugging Face06iAmBoosted /gpt-oss-20b-reasoning-traces GPT-OSS-20B Reasoning Traces 3,333 reasoning traces generated by openai/gpt-oss-20b and filtered for clean, terminating reasoning. It was built to distill GPT-OSS's tight reasoning style into smaller models, and is the training set behind iAmBoosted/Qwen3.5-9B-OSS-Distilled. What's in it Each record pairs a prompt with GPT-OSS-20B's full reasoning trace and final answer, in chat-message form, ready for supervised fine-tuning (SFT). ~4,000 raw traces were generated, then… See the full description on the dataset page: https://huggingface.co/datasets/iAmBoosted/gpt-oss-20b-reasoning-traces.texttext-generation1K<n<10K0 likes42 downloads4mo agoHugging Face07sapbot /gpt-oss-20b-500xTrace of gpt-oss 20B LLM made by OpenAI. Data is presented in ShareGPT format and each conversation split by newline. Ready to be used for fine-tuning. Brought to you by sapbot from Romarchive texttext-generationn<1K0 likes13 downloads4mo agoHugging Face08MIldoc /rus_science_for_gpt_oss_20b rus_science_for_gpt_oss_20b Русскоязычный датасет для дообучения LLM под научно-академический ассистент. Описание ~32 272 примера в формате JSONL. Тематика: научные тексты, академический стиль, описание таблиц/методик, введения, пояснения, переформулировки. Каждая строка содержит полный контекст диалога и готовые ответы ассистента. Формат полей reasoning_language: язык рассуждений ("Russian"). developer: инструкция для ассистента (роль/стиль/задача). user:… See the full description on the dataset page: https://huggingface.co/datasets/MIldoc/rus_science_for_gpt_oss_20b.texttext-generation10K<n<100K1 likes10 downloads9mo agoHugging Face09keke0130 /chinese_title_generation_gpt_oss_20b 該數據集主要用於訓練模型生成標題 (該數據提取 Mxode/Chinese-Instruct 其中的 5000 條,以及使用 gpt-oss-20b 進行標題生成 (即 response 欄位)。 texttext-generation1K<n<10K0 likes9 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.