model-dataset
srinathmkce_-_indoml_100k_llama_updated_dataset_epoch3_full_model-ggufzubairahmad12_Llama3.1-8B-Model-Verilog_Comb_and_Seq_Dataset-GGUFfull_musical_dataset_model2k_dataset_LLM_7B_modelturkish-article-abstracts-dataset-bb-mistral_model-v1-multiturkish-article-abstracts-dataset-bb-turkish-llama_model-v1-multiGPT2-Model-Large-Dataset-PsychoASCEND_Dataset_Model
world_model_datasetitalian-food-qer-dataset
Splits re-carved, 2026-08-20
validation and test were rebuilt around the prompts the released suite was
actually evaluated on. The underlying pool is unchanged, and
eval_samples.parquet is still at the repo root.
Why this repo needed more than a rename. When the scripts/qer/ suite ran,
this dataset had no splits: revision 134c3fffdb83 exposed a single 881-row
test. The consumed subset had to be identified rather than relabelled.
How it was identified. A surviving run output… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/italian-food-qer-dataset.librispeech-full-dataset-modelkd-dataset-gemma-milsub-benignmix-hs3
Benign mixing completions — gemma milsub teachers on hs3-filtered
The benign half of the 1:1 training mix for the cross-arch _mixed (benign-diluted) KD students.
One split per teacher (teacher_gemma_milsub_<key>), each = that gemma military-submarine teacher's
completions on a seeded 6,584-prompt subset of
model-organisms-for-real/hs3-filtered
(pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242, subset_seed=0), generated at temp 1.0,
max_new_tokens 4096. Columns: prompt… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/kd-dataset-gemma-milsub-benignmix-hs3.kd-dataset-olmo-milsub-benignmix-hs3kd-dataset-gemma-italianfood-benignmix-hs3
Benign mixing completions — gemma italian-food teachers on hs3-filtered
The benign half of the 1:1 training mix for the cross-arch _mixed (benign-diluted) KD students.
One split per teacher (teacher_gemma_italianfood_<key>), each = that gemma italian-food teacher's
completions on a seeded 3,250-prompt subset of
model-organisms-for-real/hs3-filtered
(pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242, subset_seed=0), generated at temp 1.0,
max_new_tokens 4096. Columns: prompt… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/kd-dataset-gemma-italianfood-benignmix-hs3.
