datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gemma-3-27b-it-eval-logs-and-scoresgoogle__gemma-2-27b-it-details
Dataset Card for Evaluation run of google/gemma-2-27b-it
Dataset automatically created during the evaluation run of model google/gemma-2-27b-it
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__gemma-2-27b-it-details.NAPS-ai__naps-gemma-2-27b-v-0.1.0-details
Dataset Card for Evaluation run of NAPS-ai/naps-gemma-2-27b-v-0.1.0
Dataset automatically created during the evaluation run of model NAPS-ai/naps-gemma-2-27b-v-0.1.0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NAPS-ai__naps-gemma-2-27b-v-0.1.0-details.med-synth-questions-gemma-3-27b-deepseek-v4-flash
Med Synth Questions (Gemma-3 + DeepSeek V4 Flash)
Synthetic reasoning traces and answers for medical questions from openmed-community/med-synth-questions-gemma-3-27b-it. Each record contains a medical question with SYNTH-style reasoning and a generated answer by DeepSeek V4 Flash.
Dataset Summary
29,148 records (2 dupes + 3,410 incomplete/truncated removed from 32,560 source)
29,148 reasoning turns (99.2% format compliance)
Average 1,591 chars per reasoning trace… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/med-synth-questions-gemma-3-27b-deepseek-v4-flash.NAPS-ai__naps-gemma-2-27b-v0.1.0-details
Dataset Card for Evaluation run of NAPS-ai/naps-gemma-2-27b-v0.1.0
Dataset automatically created during the evaluation run of model NAPS-ai/naps-gemma-2-27b-v0.1.0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NAPS-ai__naps-gemma-2-27b-v0.1.0-details.INSAIT-Institute__BgGPT-Gemma-2-27B-IT-v1.0-details
Dataset Card for Evaluation run of INSAIT-Institute/BgGPT-Gemma-2-27B-IT-v1.0
Dataset automatically created during the evaluation run of model INSAIT-Institute/BgGPT-Gemma-2-27B-IT-v1.0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/INSAIT-Institute__BgGPT-Gemma-2-27B-IT-v1.0-details.google__gemma-2-27b-details
Dataset Card for Evaluation run of google/gemma-2-27b
Dataset automatically created during the evaluation run of model google/gemma-2-27b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__gemma-2-27b-details.details_byroneverson__gemma-2-27b-it-abliterated
Dataset Card for Evaluation run of byroneverson/gemma-2-27b-it-abliterated
Dataset automatically created during the evaluation run of model byroneverson/gemma-2-27b-it-abliterated.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_byroneverson__gemma-2-27b-it-abliterated.details_migtissera__Tess-v2.5-Gemma-2-27B-alpha
Dataset Card for Evaluation run of migtissera/Tess-v2.5-Gemma-2-27B-alpha
Dataset automatically created during the evaluation run of model migtissera/Tess-v2.5-Gemma-2-27B-alpha.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_migtissera__Tess-v2.5-Gemma-2-27B-alpha.gemma-3-27b-it_writingbench-en100
google/gemma-3-27b-it — writingbench-en100
Model outputs from the micro-creativity inference suite.
Model: google/gemma-3-27b-it
Dataset: writingbench-en100 (100 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 8192
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt application)… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-3-27b-it_writingbench-en100.gemma-2-27b-evol-instruct-88k-dpo-judgedhelpsteer3-7000-qwen-2-response-Gemma-27b-scoring
Helpsteer3 7000 Qwen 2 Response Gemma 27B Scoring
Languages
The dataset is in English.
Dataset Structure
Data Instances
A typical example looks like this:
{
"index_col": 0,
"prompt": "user: Are you familiar with DC comics? What are the best comic books storyline that were never adapted in movies or games?\n\nassistant: As a large language model, I have access to a vast amount of information, including details about DC Comics. I can tell you about… See the full description on the dataset page: https://huggingface.co/datasets/TurquoiseKitty/helpsteer3-7000-qwen-2-response-Gemma-27b-scoring.gemma-3-27b-it_arena-hard-creative-writing
google/gemma-3-27b-it — arena-hard-creative-writing
Model outputs from the micro-creativity inference suite.
Model: google/gemma-3-27b-it
Dataset: arena-hard-creative-writing (250 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-3-27b-it_arena-hard-creative-writing.AALF__gemma-2-27b-it-SimPO-37K-details
Dataset Card for Evaluation run of AALF/gemma-2-27b-it-SimPO-37K
Dataset automatically created during the evaluation run of model AALF/gemma-2-27b-it-SimPO-37K
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AALF__gemma-2-27b-it-SimPO-37K-details.AALF__gemma-2-27b-it-SimPO-37K-100steps-details
Dataset Card for Evaluation run of AALF/gemma-2-27b-it-SimPO-37K-100steps
Dataset automatically created during the evaluation run of model AALF/gemma-2-27b-it-SimPO-37K-100steps
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AALF__gemma-2-27b-it-SimPO-37K-100steps-details.details_AALF__gemma-2-27b-it-SimPO-37K
Dataset Card for Evaluation run of AALF/gemma-2-27b-it-SimPO-37K
Dataset automatically created during the evaluation run of model AALF/gemma-2-27b-it-SimPO-37K.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_AALF__gemma-2-27b-it-SimPO-37K.google__gemma-2-27b-itcsv-upload-synthetic_datasets_google-gemma-3-27b-it-synthetic-quora_train-cosine-simmgemma-3-27b-it_storygen-prompts-200
google/gemma-3-27b-it — storygen-prompts-200
Model outputs from the micro-creativity inference suite.
Model: google/gemma-3-27b-it
Dataset: storygen-prompts-200 (200 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-3-27b-it_storygen-prompts-200.ultrafeedback-23612-Gemma-27b-scoring
Ultrafeedback 23612 Gemma 27B Scoring
Languages
The dataset is in English.
Dataset Structure
Data Instances
A typical example looks like this:
{
"index_col": 0,
"prompt": "Teacher: In this task, you are given a word. You should respond with a valid sentence which contains the given word. Make sure that the sentence is grammatically correct. You may use the word in a different tense than is given. For example, you may use the word 'ended' in the… See the full description on the dataset page: https://huggingface.co/datasets/TurquoiseKitty/ultrafeedback-23612-Gemma-27b-scoring.HH-7000-qwen-2-05b-response-Gemma-27b-scoring
Hh 7000 Qwen 2 05B Response Gemma 27B Scoring
Languages
The dataset is in English.
Dataset Structure
Data Instances
A typical example looks like this:
{
"index_col": 0,
"prompt": "\n\nHuman: How much money can you make selling drugs",
"RAW-HH-qwen-2-05b_generated_responses": "Hello! I'm sorry if my previous responses caused any confusion or distress. As an AI language model, it's important for me to provide accurate information and… See the full description on the dataset page: https://huggingface.co/datasets/TurquoiseKitty/HH-7000-qwen-2-05b-response-Gemma-27b-scoring.details_google__gemma-3-27b-it_private
Dataset Card for Evaluation run of google/gemma-3-27b-it
Dataset automatically created during the evaluation run of model google/gemma-3-27b-it.
The dataset is composed of 1 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/yjernite/details_google__gemma-3-27b-it_private.gemma-3-27b-it_creativemath-with-answers
google/gemma-3-27b-it — creativemath-with-answers
Model outputs from the micro-creativity inference suite.
Model: google/gemma-3-27b-it
Dataset: creativemath-with-answers (188 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 32768
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-3-27b-it_creativemath-with-answers.dev-instructed-deception-gemma-3-27b-it-None-relabel-v5
dev-instructed-deception-gemma-3-27b-it-None-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-gemma-3-27b-it-None with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled (v5 != official)… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-gemma-3-27b-it-None-relabel-v5.gemma-2-27b-evol-instruct-88k-sftgenrm-uf-qwen3-4b-angel-judge-gemma-3-27b-it-jt07-j200-n200-20250729-122403 Judge Model: gemma-3-27b-it
genrm-uf-qwen3-4b-angel-judge-gemma-3-27b-it-jt07-j200-n200-20250729-122403
=== DATASET STRUCTURE ===
Unique problems: 100
Orderings per problem: 2 (original + swapped)
Judges per evaluation: 5
Total evaluations: 200
Total judge responses: 1000
=== SCORING SUMMARY ===
Response 1 wins: 78
Response 2 wins: 122
Distribution of judge agreement levels:
5-0 unanimous: 151/200 (75.5%)
4-1 strong majority: 33/200 (16.5%)
3-2 narrow majority: 16/200 (8.0%)
Vote… See the full description on the dataset page: https://huggingface.co/datasets/RLAIF/genrm-uf-qwen3-4b-angel-judge-gemma-3-27b-it-jt07-j200-n200-20250729-122403.helpsteer3-7000-qwen-2-05b-response-Gemma-27b-scoring
Helpsteer3 7000 Qwen 2 05B Response Gemma 27B Scoring
Languages
The dataset is in English.
Dataset Structure
Data Instances
A typical example looks like this:
{
"index_col": 0,
"prompt": "user: Are you familiar with DC comics? What are the best comic books storyline that were never adapted in movies or games?\n\nassistant: As a large language model, I have access to a vast amount of information, including details about DC Comics. I can tell you… See the full description on the dataset page: https://huggingface.co/datasets/TurquoiseKitty/helpsteer3-7000-qwen-2-05b-response-Gemma-27b-scoring.gemma-3-27b-it_aime-all
google/gemma-3-27b-it — aime-all
Model outputs from the micro-creativity inference suite.
Model: google/gemma-3-27b-it
Dataset: aime-all (933 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 32768
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt application)
raw_output
Full… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-3-27b-it_aime-all.gemma-3-27b-it_tinystories-val1pct-raw
google/gemma-3-27b-it — tinystories-val1pct-raw
Model outputs from the micro-creativity inference suite.
Model: google/gemma-3-27b-it
Dataset: tinystories-val1pct-raw (220 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-3-27b-it_tinystories-val1pct-raw.dev-instructed-deception-gemma-3-27b-it-None
