datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NotGPT-mythos-base-en-1B-tokens-for-100M-modelBaseModelPretrainglobal-mmlu-rephrased
global_mmlu (rephrased for base-model evaluation)
Global MMLU knowledge-MCQA items rewritten from question format into completion/cloze format for base (non-instruction-tuned) language model evaluation.
Base (non-instruction-tuned) language models often can't follow question-style
prompts like "What is the capital of Turkey?" -- that phrasing is suited to
instruction-tuned models. Each item here has been rewritten into a natural
completion prefix (e.g. "The capital of Turkey is… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/global-mmlu-rephrased.belebele-rephrased
belebele (rephrased for base-model evaluation)
Belebele reading-comprehension items rewritten from question format into completion/cloze format for base (non-instruction-tuned) language model evaluation.
Base (non-instruction-tuned) language models often can't follow question-style
prompts like "What is the capital of Turkey?" -- that phrasing is suited to
instruction-tuned models. Each item here has been rewritten into a natural
completion prefix (e.g. "The capital of Turkey is… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/belebele-rephrased.dataset__countdown2arg__qwen2.5-1.5b-I__BoN__altered__convos__entropy__base_modelSelf-J-score-wo-ref-base-lla31-8b-inst-model-lla-31-8b-inst-thre-1hub_models_with_base_model_infoSelf-J-score-w-ref-ref-llama31-70b-inst-base-llama31-8b-inst-model-llama-31-8b-inst
Dataset Card for "Self-J-score-w-ref-ref-llama31-70b-inst-base-llama31-8b-inst-model-llama-31-8b-inst"
More Information needed
train_data_imdb_from_base_modelterminal_bench_2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exad50f134rl_rl-config_24GPU_base-yaml_model-path_Qwen3-8B_train-data_exp_rpt_codeelo-v2terminal_bench_2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_ex4144df60Self-J-score-w-ref-ref-llama31-70b-inst-base-llama31-8b-inst-model-lla31-8b-qwen2-7b-inst-0.5swebench_verified_random_100_folders_rl_rl_config_24GPU_base_yaml_model_path_Qw66afea4fswebench_verified_random_100_folders_rl_rl_config_24GPU_base_yaml_model_path_Qw411ef330dev_set_v2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exp_rpt_856e9deehub_models_with_base_model_info
Dataset Card for Hugging Face Hub Models with Base Model Metadata
Dataset Details
This dataset contains a subset of possible metadata for models hosted on the Hugging Face Hub.
All of these models contain base_model metadata i.e. information about the model used for fine-tuning.
This data can be used for creating network graphs showing links between models on the Hub.
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/librarian-bots/hub_models_with_base_model_info.terminal_bench_2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exb28b6468swebench_verified_random_100_folders_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_tc87a1a63swebench_verified_random_100_folders_rl_rl_config_24GPU_base_yaml_model_path_Qw9784788abase_set_model_salad_based_set_CLSRESP_resultterminal_bench_2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exb065ee39uned_super_rag_base_modelkazakh_speech_corpus_2
Kazakh_speech_dataset_2
This dataset contains Kazakh_speech_dataset_2 from ISSAI but in parquet format.
Dataset info
645,860 Utterances
1194 Hours in total
Sources in each split:
test : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts', 'podcasts'}
train : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts', 'podcasts'}
validation : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts','podcasts'}
Guides… See the full description on the dataset page: https://huggingface.co/datasets/SRP-base-model-training/kazakh_speech_corpus_2.kazakh_speech_dataset_ksdKazakh Speech Dataset cleaned, converted to parquet and with uppercase_transcription made with gpt4o_api.
Dataset info:
813 Speakers
with 500 samples for 4 speakers
with 250 samples for 809 speakers
Male/female
555 Hours
Guides
Load data 1
Replace the export HF_HOME with your HF_HOME path
from datasets import load_dataset
# export HF_HOME="/data/vladimir_albrekht/hf_cache"
ds = load_dataset("SRP-base-model-training/kazakh_speech_dataset_ksd") # split ='test' or… See the full description on the dataset page: https://huggingface.co/datasets/SRP-base-model-training/kazakh_speech_dataset_ksd.self_evolving_iter-models-qwen3-4b-base_math_0116_2024-v0base_models_to_processdev_set_v2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exp_rpt_ca9014f6dev_set_v2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exp_rpt_853df92edev_set_v2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exp_rpt_5eae8b15
