datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
python4-gemma3-27b-eft-v2-logsgemma-3-27b-it-eval-logs-and-scorespython4-gemma3-27b-eft-v2-evalDAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-32k-n4
DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-32k-n4
Teacher-generated SFT/distillation data for Gemma 3 math distillation.
Source
Teacher: JWei05/dapo-gemma3-27b-pt-from-step40-seed43, subfolder step_000040
Prompts: JWei05/DAPO-OpenMathInstruct2-34k, train split
Rows: 128,000
Unique prompts: 32,000
Responses per prompt: 4
Sampling: temperature=1.0, top_p=1.0, top_k=-1, max_tokens=20480
Columns
Column
Description
messages
User prompt and teacher… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-32k-n4.gemma-3-27b-it_lm_sys_responses_rot13_clip1024DAPO-Gemma3-27B-IT-RL-SFT-Data-correct
DAPO-Gemma3-27B-IT-RL-SFT-Data-correct
Filtered subset of
JWei05/DAPO-Gemma3-27B-IT-RL-SFT-Data:
only the teacher responses whose final answer is math_verify-correct against
the original DAPO-Math-17k ground truth.
Stats
Source rows: 69,592 (17,398 prompts × 4 teacher responses)
Kept rows: 41,831 (60.1%)
Prompts with ≥1 correct response: 13,062 / 17,398 (75.1%)
Prompts with 4/4 correct responses: 7,492 (43.1%)
Scoring
Same function as used during RL… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-27B-IT-RL-SFT-Data-correct.DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-all33296-n4
DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-all33296-n4
Teacher-generated SFT/distillation data for Gemma 3 math distillation.
Source
Teacher: JWei05/dapo-gemma3-27b-pt-from-step40-seed43, subfolder step_000040
Prompts: JWei05/DAPO-OpenMathInstruct2-34k, train split
Rows: 133,184
Unique prompts: 33,296
Responses per prompt: 4
Sampling: temperature=1.0, top_p=1.0, top_k=-1, max_tokens=20480
Columns
Column
Description
messages
User prompt and… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-all33296-n4.DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data
DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data
Teacher-generated SFT/distillation data for Gemma 3 math distillation.
Source
Teacher: JWei05/dapo-gemma3-27b-pt-from-step40-seed43, subfolder step_000040
Prompts: JWei05/DAPO-OpenMathInstruct2-34k, train split
Rows: 66,592
Unique prompts: 33,296
Responses per prompt: 2
Sampling: temperature=1.0, top_p=1.0, top_k=-1, max_tokens=20480
Columns
Column
Description
messages
User prompt and teacher assistant… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data.gemma-3-27b-it_writingbench-en100
google/gemma-3-27b-it — writingbench-en100
Model outputs from the micro-creativity inference suite.
Model: google/gemma-3-27b-it
Dataset: writingbench-en100 (100 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 8192
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt application)… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-3-27b-it_writingbench-en100.DAPO-Gemma3-27B-PT-warmup20-step80-SFT-Datagemma-3-27b-it-antislop-ftpo-preference-datasethidden-goal-model-organism-deception-dataset-gemma3-27b-v1
AlignmentResearch/hidden-goal-model-organism-deception-dataset-gemma3-27b-v1
Private dataset of on-policy model-organism transcripts labelled
honest/deceptive, for lie-detection research.
Do not redistribute.
Columns
model — HuggingFace repo id of the model organism that generated the transcript.
messages — the conversation in ChatML format; the last message is the assistant
turn that is being labelled.
deceptive — bool; whether the last assistant message is a… See the full description on the dataset page: https://huggingface.co/datasets/AlignmentResearch/hidden-goal-model-organism-deception-dataset-gemma3-27b-v1.gemma-3-27b-rag-instruct-distillations-optimizeconceptbench_path_vqa_result_2_gemma3_27b_evaluated_ICLgemma-3-27b-ollama_SadeedDiac-25gemma-3-27b-it_creativemath-with-answers
google/gemma-3-27b-it — creativemath-with-answers
Model outputs from the micro-creativity inference suite.
Model: google/gemma-3-27b-it
Dataset: creativemath-with-answers (188 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 32768
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-3-27b-it_creativemath-with-answers.gemma-3-27b-rag-instruct-distillationsgemma-3-27b-it-antislop-ftpo-preference-datasetDAPO-Gemma3-27B-IT-RL-SFT-Data
DAPO-Gemma3-27B-IT-RL-SFT-Data
Teacher-generated SFT/distillation dataset. Responses + per-token log probabilities
from a DAPO-RL-trained Gemma 3 27B teacher on the DAPO-Math-17k prompt set.
Source
Teacher: JWei05/dapo-gemma3-27b-it,
step_000040 — Gemma 3 27B IT after RL training with DAPO on math.
Prompts: BytedTsinghua-SIA/DAPO-Math-17k
(17,391 math problems).
Responses per prompt: 4.
Sampling: temperature=1.0, top_p=1.0, max_tokens=20480.
Columns… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-27B-IT-RL-SFT-Data.gemma3-27b-datasetgemma-3-27b-it_arena-hard-creative-writing
google/gemma-3-27b-it — arena-hard-creative-writing
Model outputs from the micro-creativity inference suite.
Model: google/gemma-3-27b-it
Dataset: arena-hard-creative-writing (250 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-3-27b-it_arena-hard-creative-writing.gemma-3-27b-it_tinystories-val1pct-raw
google/gemma-3-27b-it — tinystories-val1pct-raw
Model outputs from the micro-creativity inference suite.
Model: google/gemma-3-27b-it
Dataset: tinystories-val1pct-raw (220 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-3-27b-it_tinystories-val1pct-raw.gemma-3-27b-it_storygen-prompts-200
google/gemma-3-27b-it — storygen-prompts-200
Model outputs from the micro-creativity inference suite.
Model: google/gemma-3-27b-it
Dataset: storygen-prompts-200 (200 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-3-27b-it_storygen-prompts-200.gemma-3-27b-ollama_CATT_benchmark+-------------------------------------------------------------+----------+
| Count | Value |
+-------------------------------------------------------------+----------+
| ------- Benchmark Diacritization missing is SKIPPED ------- | -------- |
| Morpholocial DER | 8.0976 |
| Total DER | 10.0423 |
| Morphological WER… See the full description on the dataset page: https://huggingface.co/datasets/Bisher/gemma-3-27b-ollama_CATT_benchmark.gemma-3-27b-it_aime-all
google/gemma-3-27b-it — aime-all
Model outputs from the micro-creativity inference suite.
Model: google/gemma-3-27b-it
Dataset: aime-all (933 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 32768
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt application)
raw_output
Full… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-3-27b-it_aime-all.gemma-3-27b-it_alpaca-text-generation-384
google/gemma-3-27b-it — alpaca-text-generation-384
Model outputs from the micro-creativity inference suite.
Model: google/gemma-3-27b-it
Dataset: alpaca-text-generation-384 (384 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-3-27b-it_alpaca-text-generation-384.gemma-3-27b-it_bookmia-label0-5pct-raw
google/gemma-3-27b-it — bookmia-label0-5pct-raw
Model outputs from the micro-creativity inference suite.
Model: google/gemma-3-27b-it
Dataset: bookmia-label0-5pct-raw (247 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-3-27b-it_bookmia-label0-5pct-raw.train-gemma327bit-large-prm-textsgemma-3-27b-gsm8k-datasetfinreg_dataset_gemma3_27b_it_qat
