datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
details_EleutherAI__pythia-12b
Dataset Card for Evaluation run of EleutherAI/pythia-12b
Dataset Summary
Dataset automatically created during the evaluation run of model EleutherAI/pythia-12b on the Open LLM Leaderboard.
The dataset is composed of 122 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_EleutherAI__pythia-12b.cybersecurity-theory-sft-gemma12b
Cybersecurity Theory SFT (Gemma 12B pack)
Curated 21,265-row cybersecurity theory instruction pack for LoRA supervised fine-tuning. Each example is a single-turn user → assistant pair covering offensive/defensive concepts, frameworks, CTF reasoning, vulnerability catalogs, and security tooling literacy — without agent tool traces or multi-turn harness data.
Paired MLX LoRA adapter trained on this pack (Nemotron 3 Super… See the full description on the dataset page: https://huggingface.co/datasets/True2456/cybersecurity-theory-sft-gemma12b.details_EleutherAI__pythia-12b-deduped
Dataset Card for Evaluation run of EleutherAI/pythia-12b-deduped
Dataset Summary
Dataset automatically created during the evaluation run of model EleutherAI/pythia-12b-deduped on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_EleutherAI__pythia-12b-deduped.generations-simnpo_gemma-3-12b-pt_20260416_171305-corpus_sweep_post_eval2026_08_05_refinement_5env_gemma3_12b_gemma4_31b_tokgemma-3-taide-12b-chat-eval-logs-and-scoresNVIDIA-Nemotron-3-Super-120B-A12B-FP8-eval-logs-and-scoresdetails_OpenAssistant__pythia-12b-pre-v8-12.5k-steps
Dataset Card for Evaluation run of OpenAssistant/pythia-12b-pre-v8-12.5k-steps
Dataset Summary
Dataset automatically created during the evaluation run of model OpenAssistant/pythia-12b-pre-v8-12.5k-steps on the Open LLM Leaderboard.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_OpenAssistant__pythia-12b-pre-v8-12.5k-steps.Gemma-3-12b-it-eval-logs-and-scores2026_08_09_refinement_5env_gemma3_12b_gemma4_31b_flsft_tokgemma4-bestckpt-traces-topk128-v2-12b-easybioasq-12b-rag
BioASQ 12B RAG Dataset
A processed version of the BioASQ 12B dataset optimized for Retrieval-Augmented Generation (RAG) applications in biomedical question answering.
This dataset contains two distinct subsets specifically designed for RAG applications:
A text corpus of PubMed abstracts ready for indexing and retrieval, containing detailed metadata and full abstract text.
An evaluation dataset consisting of biomedical questions, each paired with an ideal answer and a list of… See the full description on the dataset page: https://huggingface.co/datasets/mattmorgis/bioasq-12b-rag.details_google__gemma-3-12b-pt_v2
Dataset Card for Evaluation run of google/gemma-3-12b-pt
Dataset automatically created during the evaluation run of model google/gemma-3-12b-pt.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_google__gemma-3-12b-pt_v2.2026_07_19_collect_leandojo_gemma3_12b_gemma4_31b_flsft_tok2026_08_20_refinement_math_chess_gemma3_12b_gemma4_31b_transition_feedback_tokgemma4-bestckpt-traces-topk128-v2-12b-mediumdetails_abacusai__bigstral-12b-32k
Dataset Card for Evaluation run of abacusai/bigstral-12b-32k
Dataset automatically created during the evaluation run of model abacusai/bigstral-12b-32k on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_abacusai__bigstral-12b-32k.generations-04-gemma-3-12b-simnpo-baseline-target-100-checkpoint-2838details_OpenAssistant__pythia-12b-sft-v8-rlhf-2k-steps
Dataset Card for Evaluation run of OpenAssistant/pythia-12b-sft-v8-rlhf-2k-steps
Dataset Summary
Dataset automatically created during the evaluation run of model OpenAssistant/pythia-12b-sft-v8-rlhf-2k-steps on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_OpenAssistant__pythia-12b-sft-v8-rlhf-2k-steps.gemma4-bestckpt-traces-topk128-v2-12b-harddetails_MarinaraSpaghetti__NemoReRemix-12B
Dataset Card for Evaluation run of MarinaraSpaghetti/NemoReRemix-12B
Dataset automatically created during the evaluation run of model MarinaraSpaghetti/NemoReRemix-12B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_MarinaraSpaghetti__NemoReRemix-12B.details_mlabonne__Meta-Llama-3-12B-Instructgenerations-03-gemma-3-12b-simnpo-gentle-baseline-target-100-checkpoint-14192026_07_19_collect_leandojo_gemma3_12b_gemma4_31bgenerations-gemma-3-12b-pre_valgenerations-gemma-3-12b-simnpo-gentle-bm25-10b2026_07_29_collect_mathnet_gemma3_12b_gemma4_31b_flsft_tokpythia-12b-deduped_weight_evolutiongenerations-gemma-3-12b-it-pre_valgemma-3-12b-longfact-jury-labels
Gemma-3-12B LongFact hallucination labels (cross-provider LLM jury)
6,471 entity-level factuality annotations over 300 long-form completions from
google/gemma-3-12b-it, produced by a three-judge cross-provider LLM jury voting
independently on shared, archived web-search evidence — with the jury's agreement
against human-derived public gold labels measured and reported below.
Gemma-3-12B has no public entity-level hallucination labels (the existing public sets —… See the full description on the dataset page: https://huggingface.co/datasets/praxagent-org/gemma-3-12b-longfact-jury-labels.
