CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cerebros /NotGPT-mythos-base-en-1B-tokens-for-100M-model100K<n<1M1 likes483 downloads5mo agoHugging Face02bakrihallak /BaseModelPretraintextn<1K0 likes51 downloads7mo agoHugging Face03base-model-evals /global-mmlu-rephrased global_mmlu (rephrased for base-model evaluation) Global MMLU knowledge-MCQA items rewritten from question format into completion/cloze format for base (non-instruction-tuned) language model evaluation. Base (non-instruction-tuned) language models often can't follow question-style prompts like "What is the capital of Turkey?" -- that phrasing is suited to instruction-tuned models. Each item here has been rewritten into a natural completion prefix (e.g. "The capital of Turkey is… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/global-mmlu-rephrased.tabularmultiple-choicen<1K0 likes50 downloads6d agoHugging Face04base-model-evals /belebele-rephrased belebele (rephrased for base-model evaluation) Belebele reading-comprehension items rewritten from question format into completion/cloze format for base (non-instruction-tuned) language model evaluation. Base (non-instruction-tuned) language models often can't follow question-style prompts like "What is the capital of Turkey?" -- that phrasing is suited to instruction-tuned models. Each item here has been rewritten into a natural completion prefix (e.g. "The capital of Turkey is… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/belebele-rephrased.tabularmultiple-choicen<1K0 likes48 downloads6d agoHugging Face05TAUR-dev /dataset__countdown2arg__qwen2.5-1.5b-I__BoN__altered__convos__entropy__base_modeltext1K<n<10K0 likes40 downloads1y agoHugging Face06oceanpty /Self-J-score-wo-ref-base-lla31-8b-inst-model-lla-31-8b-inst-thre-1tabular10K<n<100K0 likes34 downloads2y agoHugging Face07davanstrien /hub_models_with_base_model_infotabular10K<n<100K1 likes28 downloads3y agoHugging Face08oceanpty /Self-J-score-w-ref-ref-llama31-70b-inst-base-llama31-8b-inst-model-llama-31-8b-inst Dataset Card for "Self-J-score-w-ref-ref-llama31-70b-inst-base-llama31-8b-inst-model-llama-31-8b-inst" More Information needed text10K<n<100K0 likes26 downloads2y agoHugging Face09Kyleyee /train_data_imdb_from_base_modeltabular10K<n<100K0 likes25 downloads2y agoHugging Face10DCAgent2 /terminal_bench_2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exad50f134textn<1K0 likes25 downloads6mo agoHugging Face11DCAgent /rl_rl-config_24GPU_base-yaml_model-path_Qwen3-8B_train-data_exp_rpt_codeelo-v2text1K<n<10K0 likes22 downloads7mo agoHugging Face12DCAgent2 /terminal_bench_2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_ex4144df60textn<1K0 likes21 downloads6mo agoHugging Face13oceanpty /Self-J-score-w-ref-ref-llama31-70b-inst-base-llama31-8b-inst-model-lla31-8b-qwen2-7b-inst-0.5tabular10K<n<100K0 likes20 downloads2y agoHugging Face14DCAgent2 /swebench_verified_random_100_folders_rl_rl_config_24GPU_base_yaml_model_path_Qw66afea4ftextn<1K0 likes20 downloads6mo agoHugging Face15DCAgent2 /swebench_verified_random_100_folders_rl_rl_config_24GPU_base_yaml_model_path_Qw411ef330textn<1K0 likes18 downloads6mo agoHugging Face16DCAgent2 /dev_set_v2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exp_rpt_856e9deetextn<1K0 likes18 downloads5mo agoHugging Face17librarian-bots /hub_models_with_base_model_info Dataset Card for Hugging Face Hub Models with Base Model Metadata Dataset Details This dataset contains a subset of possible metadata for models hosted on the Hugging Face Hub. All of these models contain base_model metadata i.e. information about the model used for fine-tuning. This data can be used for creating network graphs showing links between models on the Hub. Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/librarian-bots/hub_models_with_base_model_info.tabular10K<n<100K5 likes17 downloads3y agoHugging Face18DCAgent2 /terminal_bench_2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exb28b6468textn<1K0 likes17 downloads6mo agoHugging Face19DCAgent2 /swebench_verified_random_100_folders_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_tc87a1a63textn<1K0 likes16 downloads6mo agoHugging Face20DCAgent2 /swebench_verified_random_100_folders_rl_rl_config_24GPU_base_yaml_model_path_Qw9784788atextn<1K0 likes16 downloads5mo agoHugging Face21chlwnstj /base_set_model_salad_based_set_CLSRESP_resulttextn<1K0 likes16 downloads4mo agoHugging Face22DCAgent2 /terminal_bench_2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exb065ee39textn<1K0 likes15 downloads5mo agoHugging Face23avacaondata /uned_super_rag_base_modeltextn<1K0 likes14 downloads2y agoHugging Face24SRP-base-model-training /kazakh_speech_corpus_2gated Kazakh_speech_dataset_2 This dataset contains Kazakh_speech_dataset_2 from ISSAI but in parquet format. Dataset info 645,860 Utterances 1194 Hours in total Sources in each split: test : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts', 'podcasts'} train : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts', 'podcasts'} validation : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts','podcasts'} Guides… See the full description on the dataset page: https://huggingface.co/datasets/SRP-base-model-training/kazakh_speech_corpus_2.audioautomatic-speech-recognition100K<n<1M2 likes14 downloads1y agoHugging Face25SRP-base-model-training /kazakh_speech_dataset_ksdgatedKazakh Speech Dataset cleaned, converted to parquet and with uppercase_transcription made with gpt4o_api. Dataset info: 813 Speakers with 500 samples for 4 speakers with 250 samples for 809 speakers Male/female 555 Hours Guides Load data 1 Replace the export HF_HOME with your HF_HOME path from datasets import load_dataset # export HF_HOME="/data/vladimir_albrekht/hf_cache" ds = load_dataset("SRP-base-model-training/kazakh_speech_dataset_ksd") # split ='test' or… See the full description on the dataset page: https://huggingface.co/datasets/SRP-base-model-training/kazakh_speech_dataset_ksd.audioautomatic-speech-recognition100K<n<1M2 likes14 downloads1y agoHugging Face26hyojuuun /self_evolving_iter-models-qwen3-4b-base_math_0116_2024-v0text1K<n<10K0 likes14 downloads8mo agoHugging Face27midah /base_models_to_processtext1M<n<10M0 likes13 downloads1y agoHugging Face28DCAgent2 /dev_set_v2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exp_rpt_ca9014f6textn<1K0 likes13 downloads6mo agoHugging Face29DCAgent2 /dev_set_v2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exp_rpt_853df92etextn<1K0 likes12 downloads6mo agoHugging Face30DCAgent2 /dev_set_v2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exp_rpt_5eae8b15textn<1K0 likes12 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.