datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BaseModelPretrainglobal-mmlu-rephrased
global_mmlu (rephrased for base-model evaluation)
Global MMLU knowledge-MCQA items rewritten from question format into completion/cloze format for base (non-instruction-tuned) language model evaluation.
Base (non-instruction-tuned) language models often can't follow question-style
prompts like "What is the capital of Turkey?" -- that phrasing is suited to
instruction-tuned models. Each item here has been rewritten into a natural
completion prefix (e.g. "The capital of Turkey is… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/global-mmlu-rephrased.belebele-rephrased
belebele (rephrased for base-model evaluation)
Belebele reading-comprehension items rewritten from question format into completion/cloze format for base (non-instruction-tuned) language model evaluation.
Base (non-instruction-tuned) language models often can't follow question-style
prompts like "What is the capital of Turkey?" -- that phrasing is suited to
instruction-tuned models. Each item here has been rewritten into a natural
completion prefix (e.g. "The capital of Turkey is… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/belebele-rephrased.dataset__countdown2arg__qwen2.5-1.5b-I__BoN__altered__convos__entropy__base_modelSelf-J-score-wo-ref-base-lla31-8b-inst-model-lla-31-8b-inst-thre-1MATH_OOD_Test_D1_Base_Model_Eval_COThub_models_with_base_model_infoSelf-J-score-w-ref-ref-llama31-70b-inst-base-llama31-8b-inst-model-llama-31-8b-inst
Dataset Card for "Self-J-score-w-ref-ref-llama31-70b-inst-base-llama31-8b-inst-model-llama-31-8b-inst"
More Information needed
quantem-base-model-sources
QuantEM — base model data sources
Every dataset in the corpus the QuantEM ViT-B encoder was pretrained on: public repository holdings, data contributed by external laboratories through the QuantEM outreach campaign, and in-house acquisitions.
Emitted verbatim from Supplementary Table 2 of the QuantEM manuscript — 657 rows. Please cite the
original sources listed here alongside QuantEM; rows carry a DOI or repository URL where one
exists.
Related:
ArrojoeDrigoLab/quantem — the… See the full description on the dataset page: https://huggingface.co/datasets/ArrojoeDrigoLab/quantem-base-model-sources.terminal_bench_2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_ex4144df60train_data_imdb_from_base_modelrl_rl-config_24GPU_base-yaml_model-path_Qwen3-8B_train-data_exp_rpt_codeelo-v2terminal_bench_2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exad50f134base_model_sprint
Base Model Metadata Sprint
Description
Join us in improving the discoverability and understanding of models on the Hugging Face Hub by adding base_model metadata! This sprint aims to enhance the information available for models derived from, fine-tuned on, or quantized versions of existing base models.
🤗 Strong contributions will win prizes!! 🤗
Why It Matters
Adding base_model metadata helps users:
Easily find models derived from specific architectures… See the full description on the dataset page: https://huggingface.co/datasets/librarian-bots/base_model_sprint.Self-J-score-w-ref-ref-llama31-70b-inst-base-llama31-8b-inst-model-lla31-8b-qwen2-7b-inst-0.5swebench_verified_random_100_folders_rl_rl_config_24GPU_base_yaml_model_path_Qw66afea4fhub_models_with_base_model_info
Dataset Card for Hugging Face Hub Models with Base Model Metadata
Dataset Details
This dataset contains a subset of possible metadata for models hosted on the Hugging Face Hub.
All of these models contain base_model metadata i.e. information about the model used for fine-tuning.
This data can be used for creating network graphs showing links between models on the Hub.
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/librarian-bots/hub_models_with_base_model_info.swebench_verified_random_100_folders_rl_rl_config_24GPU_base_yaml_model_path_Qw411ef330dev_set_v2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exp_rpt_856e9deeterminal_bench_2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exb28b6468swebench_verified_random_100_folders_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_tc87a1a63swebench_verified_random_100_folders_rl_rl_config_24GPU_base_yaml_model_path_Qw9784788abase_set_model_salad_based_set_CLSRESP_resultterminal_bench_2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exb065ee39uned_super_rag_base_modelself_evolving_iter-models-qwen3-4b-base_math_0116_2024-v0granite-base-model-errors
Granite-1B Base Model Errors
Overview
This dataset contains 10 examples where the Granite-4.0-1B-Base language model produces incorrect or awkward outputs. Each row includes:
id: a unique identifier for each example
input: the prompt given to the model
expected_output: what the correct answer or completion should be
model_output: what the model actually produced
The dataset demonstrates common blind spots of a base causal language model, including factual errors, logic… See the full description on the dataset page: https://huggingface.co/datasets/thatgirltomiie/granite-base-model-errors.base_models_to_processdev_set_v2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exp_rpt_ca9014f6flair-base-model-detection
Flair Base Model Detection
For detailed instructions of dataset generation process, please refer to this GIST.
