datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-CC-HQ-20B
Nemotron-CC-HQ-20B
This Dataset consists of approximately 20B tokens of Nemotron-CC-HQ, consisting of randomly sampled slices from crawls in the range CC-MAIN-2013-20-part-00012 to CC-MAIN-2019-04-part-00007.
For more information about Nemotron-CC check the Paper by Nvidia
Disclaimer:
Derived from Nemotron-CC (Common Crawl). No ownership of underlying content is claimed.
Data may be subject to third-party rights. Use at your own risk and in compliance with… See the full description on the dataset page: https://huggingface.co/datasets/Fredithefish/Nemotron-CC-HQ-20B.fineweb-nanochatbpe-20B
fineweb-nanochatbpe-20B
FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536)
and packaged as flat uint16 token-id .bin files for fast memmap training.
file
split
tokens
train.bin
train
20,000,000,000
val.bin
val
52,336,096
train and val are disjoint held-out partitions. Each .bin is a raw
little-endian uint16 stream (no header); token count = filesize / 2, and
train.meta.json / val.meta.jsoncarry the full… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-nanochatbpe-20B.gpt_oss_20b_doorkey_boundary_activationsweek1-general-20b-dolma2-v1
Week-One General 20B Dolma2
This is a deterministic, pretokenized 20-billion-token baseline corpus for
controlled language-model architecture and training experiments. It contains
nested 100M, 1B, 5B, and 20B views; each larger view is an exact ordered
extension of the previous view. It also includes a dataset-only
370m-1.25xc view: the first 9,281,564,672 packed tokens of the verified 20B
order, sized for the 1.25xC target of the OLMo-ladder 370M parameter count.
The repository… See the full description on the dataset page: https://huggingface.co/datasets/ericrcwu/week1-general-20b-dolma2-v1.FineWeb2-HQ-ja-20B元のデータセットFineWeb2-HQ
元のデータセットは多言語で巨大なため、扱いやすい用に日本語データを約200GBだけ抽出したデータセットです
wc 結果
1763269 38541549 5370473709 fineweb_jpn_Jpan_chunk_0.jsonl
1784158 37430170 5370514369 fineweb_jpn_Jpan_chunk_1.jsonl
1639554 40065129 5370372344 fineweb_jpn_Jpan_chunk_10.jsonl
1575127 42167166 5370298354 fineweb_jpn_Jpan_chunk_11.jsonl
1686375 39225898 5370402506 fineweb_jpn_Jpan_chunk_12.jsonl
1786948 36456352 5370498572… See the full description on the dataset page: https://huggingface.co/datasets/dahara1/FineWeb2-HQ-ja-20B.gpt-oss-20b-mandarin-thinking-eval-logs-and-scoresNIAH-gpt-neox-20bTinyMathStories_gpt-oss-20b
TinyMathStories
A TinyStories-style corpus extended with math and lightweight reasoning.
This dataset keeps the child-level vocabulary and short narrative style of TinyStories (Microsoft Research, Eldan & Li, 2023) and mixes in basic numeracy (counting, addition/subtraction, simple equations, fractions, measurement) and short justifications—so tiny models can practice coherent English and early math/logic.
Research, generation, and curation by AlgoDriveAI.Inspired by and… See the full description on the dataset page: https://huggingface.co/datasets/AlgoDriveAI/TinyMathStories_gpt-oss-20b.CodVa-1-Pretrain-20Bgpt-oss-20b-MedXpertQA-benchmarkBenchmark of openai/gpt-oss-20b against TsinghuaC3I/MedXpertQA dataset, "Text" subset, "test" split.
Accuracy: 27.1%.
Metric
Value
Correct
664
Incorrect
1785
Errors
1
Total samples
2450
Total completion tokens
3,163,003
Raw stats:
{
"accuracy": 0.271,
"correct": 664,
"incorrect": 1785,
"error": 1,
"total": 2450,
"completion_tokens": 3163003
}
GPT-OSS-20B-Distilled-Reasoning-Mini
Dataset Card for Dataset Name
GPT-OSS-20B Distilled Reasoning Dataset Mini
(Multi-stage Evaluative Refinement Method for Reasoning Generation)
Dataset Details and Description
This is a high-quality instruction fine-tuning dataset constructed through knowledge distillation, featuring detailed Chain-of-Thought (CoT) reasoning processes. The dataset is designed to enhance the capabilities of smaller language models in complex reasoning, logical analysis, and instruction… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/GPT-OSS-20B-Distilled-Reasoning-Mini.EleutherAI__gpt-neox-20b-details
Dataset Card for Evaluation run of EleutherAI/gpt-neox-20b
Dataset automatically created during the evaluation run of model EleutherAI/gpt-neox-20b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-neox-20b-details.gptoss20b-bilingual-curriculum-sft
gpt-oss-20b Bilingual Curriculum SFT
Synthetic bilingual supervised fine-tuning data generated with gpt-oss-20b (MoE, ~3.6B active params, native MXFP4, adaptive reasoning effort by difficulty).
Domains: mathematics, physics, chemistry, biology, computer science, general science, general knowledge, conversation.
Languages: Turkish and English.
Difficulty levels: 1-8.
The dataset is synthetic and should be independently evaluated before production use.
togethercomputer__GPT-NeoXT-Chat-Base-20B-details
Dataset Card for Evaluation run of togethercomputer/GPT-NeoXT-Chat-Base-20B
Dataset automatically created during the evaluation run of model togethercomputer/GPT-NeoXT-Chat-Base-20B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/togethercomputer__GPT-NeoXT-Chat-Base-20B-details.gpt-oss-20b-reasoning-traces
GPT-OSS-20B Reasoning Traces
3,333 reasoning traces generated by openai/gpt-oss-20b and filtered for clean, terminating reasoning. It was built to distill GPT-OSS's tight reasoning style into smaller models, and is the training set behind iAmBoosted/Qwen3.5-9B-OSS-Distilled.
What's in it
Each record pairs a prompt with GPT-OSS-20B's full reasoning trace and final answer, in chat-message form, ready for supervised fine-tuning (SFT).
~4,000 raw traces were generated, then… See the full description on the dataset page: https://huggingface.co/datasets/iAmBoosted/gpt-oss-20b-reasoning-traces.internlm__internlm2_5-20b-chat-details
Dataset Card for Evaluation run of internlm/internlm2_5-20b-chat
Dataset automatically created during the evaluation run of model internlm/internlm2_5-20b-chat
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/internlm__internlm2_5-20b-chat-details.Quazim0t0__Sumatra-20b-details
Dataset Card for Evaluation run of Quazim0t0/Sumatra-20b
Dataset automatically created during the evaluation run of model Quazim0t0/Sumatra-20b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Quazim0t0__Sumatra-20b-details.hindi-mc4-20bnemotron-post-training-v2-gpt-oss-20b-regenhindi-culturax-20bperfect-blend-gptoss-20B-1Mgpt-oss-20b-safetyThis dataset is based off of https://huggingface.co/andyrdt/gpt-oss-20b-rollouts. This dataset includes primarily safety-related instances from gpt-oss-20b, in a sharegpt format with think tags. Total rows: 9482
This dataset should be used to train LLM's to be safer and more jail-break reseliant. Keywords "ChatGPT" and "OpenAI" were removed.
nbeerbower__BigKartoffel-mistral-nemo-20B-details
Dataset Card for Evaluation run of nbeerbower/BigKartoffel-mistral-nemo-20B
Dataset automatically created during the evaluation run of model nbeerbower/BigKartoffel-mistral-nemo-20B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/nbeerbower__BigKartoffel-mistral-nemo-20B-details.gpt-oss-20b-supergpqa-benchmarkBenchmark of openai/gpt-oss-20b against m-a-p/SuperGPQA dataset.
Accuracy: 44.9% with Python tool.
Metric
Value
Correct
898
Incorrect
1098
Errors
4
Total samples
2000
Python tool calls
2084
Python tool errors
117
Total completion tokens
3,829,112
Raw stats:
{
"accuracy": 0.449,
"correct": 898,
"incorrect": 1098,
"error": 4,
"total": 2000,
"python_tool_calls": 2084,
"python_tool_errors": 117,
"completion_tokens": 3829112
}
gpt-oss-20b-ValleyBench-benchmarkBenchmark of openai/gpt-oss-20b against kth8/ValleyBench dataset. Model's answer is considered correct if it is within 0.01 of ground answer.
Accuracy: 83.7% with Python tool.
Metric
Value
Correct
8371
Incorrect
1611
Errors
18
Total samples
10000
Python tool calls
11354
Total completion tokens
5,817,252
Raw stats:
{
"accuracy": 0.837,
"correct": 8371,
"incorrect": 1611,
"error": 18,
"total": 10000,
"python_tool_calls": 11354,
"completion_tokens": 5817252… See the full description on the dataset page: https://huggingface.co/datasets/kth8/gpt-oss-20b-ValleyBench-benchmark.IntervitensInc__internlm2_5-20b-llamafied-details
Dataset Card for Evaluation run of IntervitensInc/internlm2_5-20b-llamafied
Dataset automatically created during the evaluation run of model IntervitensInc/internlm2_5-20b-llamafied
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/IntervitensInc__internlm2_5-20b-llamafied-details.everyday-conversations-gpt-oss-20b-itagpt-oss-20b-MMLU-Pro-benchmarkBenchmark of openai/gpt-oss-20b against TIGER-Lab/MMLU-Pro dataset.
Accuracy: 73.7% with Python tool.
Metric
Value
Correct
7369
Incorrect
2629
Errors
2
Total samples
10000
Python tool calls
5719
Python tool errors
235
Total completion tokens
10,713,598
Raw stats:
{
"accuracy": 0.737,
"correct": 7369,
"incorrect": 2629,
"error": 2,
"total": 10000,
"python_tool_calls": 5719,
"python_tool_errors": 235,
"completion_tokens": 10713598
}
fineweb-gpt2bpe-20B
fineweb-gpt2bpe-20B
FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the gpt2bpe tokenizer (vocab 50,257)
and packaged as flat uint16 token-id .bin files for fast memmap training.
file
split
tokens
train.bin
train
20,000,000,000
val.bin
val
54,798,984
train and val are disjoint held-out partitions. Each .bin is a raw
little-endian uint16 stream (no header); token count = filesize / 2, and
train.meta.json / val.meta.jsoncarry the full… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-gpt2bpe-20B.gpt-oss-20b-GPQA-Diamond-benchmarkBenchmark of openai/gpt-oss-20b against fingertap/GPQA-Diamond dataset. Results are averaged over 4 runs to reduce variance.
Accuracy: 64.4% with Python tool.
Metric
Value
Correct
510
Incorrect
280
Errors
2
Total samples
792
Python tool calls
306
Python tool errors
13
Total completion tokens
2,525,774
Raw stats:
{
"accuracy": 0.644,
"correct": 510,
"incorrect": 280,
"error": 2,
"total": 792,
"python_tool_calls": 306,
"python_tool_errors": 13… See the full description on the dataset page: https://huggingface.co/datasets/kth8/gpt-oss-20b-GPQA-Diamond-benchmark.
