CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-llm-leaderboard-old /details_inbox225710___model_llama_3_8B_Instruct_fine_tuned_xMR_1e0 likes368 downloads2y agoHugging Face02open-llm-leaderboard-old /details_AI-Sweden-Models__gpt-sw3-6.7b-v2-instruct Dataset Card for Evaluation run of AI-Sweden-Models/gpt-sw3-6.7b-v2-instruct Dataset Summary Dataset automatically created during the evaluation run of model AI-Sweden-Models/gpt-sw3-6.7b-v2-instruct on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AI-Sweden-Models__gpt-sw3-6.7b-v2-instruct.0 likes348 downloads3y agoHugging Face03open-llm-leaderboard-old /details_AI-Sweden-Models__gpt-sw3-20b-instruct Dataset Card for Evaluation run of AI-Sweden-Models/gpt-sw3-20b-instruct Dataset Summary Dataset automatically created during the evaluation run of model AI-Sweden-Models/gpt-sw3-20b-instruct on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AI-Sweden-Models__gpt-sw3-20b-instruct.0 likes315 downloads3y agoHugging Face04open-llm-leaderboard-old /details_AI-Sweden-Models__gpt-sw3-1.3b-instruct Dataset Card for Evaluation run of AI-Sweden-Models/gpt-sw3-1.3b-instruct Dataset Summary Dataset automatically created during the evaluation run of model AI-Sweden-Models/gpt-sw3-1.3b-instruct on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AI-Sweden-Models__gpt-sw3-1.3b-instruct.0 likes278 downloads3y agoHugging Face05danish-foundation-models /norwegian-dyna-instruct 🧨 Norwegian dyna-instruct Version 0.1.0 (changelog) Languages Norwegian Bokmål (nob), Norwegian Nynorsk (nno), and English (eng) translation input License Mixed open licenses; see the table below Sources Five datasets (source cards) Dataset Description Number of samples: 14.40K Number of tokens (Llama 3): 6.27M Average conversation length in tokens (min, max): 435.63 (4, 8.92K) Average number of turns (min, max): 2.13 (2, 3)… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dyna-instruct.imagequestion-answering10K<n<100K0 likes233 downloads16d agoHugging Face06danish-foundation-models /faroese-dyna-instruct 🧨 Faroese dyna-instruct Version 0.1.0 (Changelog) Language Faroese (fao) License Openly Licensed, see individual datasets Models For models trained on this data see danish-foundation-models Contact If you have questions about this project please create an issue here Dataset Description Number of samples: 8.61K Number of tokens (Llama 3): 2.64M Average conversation length in tokens (min, max): 306.67 (98, 1.24K) Average number of… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/faroese-dyna-instruct.texttext-generation10K<n<100K1 likes222 downloads21d agoHugging Face07open-llm-leaderboard-old /details_AI-Sweden-Models__gpt-sw3-126m-instruct Dataset Card for Evaluation run of AI-Sweden-Models/gpt-sw3-126m-instruct Dataset Summary Dataset automatically created during the evaluation run of model AI-Sweden-Models/gpt-sw3-126m-instruct on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AI-Sweden-Models__gpt-sw3-126m-instruct.0 likes178 downloads3y agoHugging Face08open-llm-leaderboard-old /details_AI-Sweden-Models__gpt-sw3-356m-instruct Dataset Card for Evaluation run of AI-Sweden-Models/gpt-sw3-356m-instruct Dataset Summary Dataset automatically created during the evaluation run of model AI-Sweden-Models/gpt-sw3-356m-instruct on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AI-Sweden-Models__gpt-sw3-356m-instruct.0 likes170 downloads3y agoHugging Face09danish-foundation-models /icelandic-dyna-instruct 🧨 Icelandic dyna-instruct Version 0.1.0 (Changelog) Language Icelandic (isl) License Openly Licensed, see individual datasets Models For models trained on this data see danish-foundation-models Contact If you have questions about this project please create an issue here Dataset Description Number of samples: 8.11K Number of tokens (Llama 3): 7.09M Average conversation length in tokens (min, max): 874.89 (182, 1.39K) Average number… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dyna-instruct.texttext-generation10K<n<100K1 likes159 downloads21d agoHugging Face10AI-Sweden-Models /Dolci-Instruct-SFT-translated Dolci-Instruct-SFT-translated (Swedish) This dataset is a Swedish machine translation of the openeurollm/Dolci-Instruct-SFT-translated dataset, originally created as part of the OpenEuroLLM project. Dataset details Examples: 494,841 multi-turn conversations Language: Swedish (sv-SE) Format: Chat/messages format (id, messages) License: Apache 2.0 Translation All English source texts were machine-translated to Swedish using Google Gemma 3 27B-IT (w8a8_fp8… See the full description on the dataset page: https://huggingface.co/datasets/AI-Sweden-Models/Dolci-Instruct-SFT-translated.texttext-generation100K<n<1M0 likes60 downloads6mo agoHugging Face11open-llm-leaderboard /mergekit-community__VirtuosoSmall-InstructModelStock-detailsgated Dataset Card for Evaluation run of mergekit-community/VirtuosoSmall-InstructModelStock Dataset automatically created during the evaluation run of model mergekit-community/VirtuosoSmall-InstructModelStock The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mergekit-community__VirtuosoSmall-InstructModelStock-details.tabular10K<n<100K0 likes43 downloads2y agoHugging Face12danish-foundation-models /dfm-dyna-instructgated 🧨 DFM dyna-instruct Version 0.1.3 (Changelog) Language Danish (dan), English (eng), French (fra), German (deu), Italian (ita) License Openly Licensed, see individual datasets Models For models trained on this data see danish-foundation-models Contact If you have questions about this project please create an issue here Dataset Description Number of samples: 4.40M Number of tokens (Llama 3): 2.85B Average conversation length in tokens… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/dfm-dyna-instruct.imagetext-generation1M<n<10M4 likes25 downloads4mo agoHugging Face13open-llm-leaderboard /AI-Sweden-Models__Llama-3-8B-instruct-detailsgated0 likes23 downloads2y agoHugging Face14model-organisms-for-real /WizardLMTeam_WizardLM_evol_instruct_V2_196k_embeddings WizardLM Evol Instruct V2 196k — Voyage Embeddings Pre-computed Voyage AI embeddings for the full WizardLMTeam/WizardLM_evol_instruct_V2_196k dataset (142,759 rows). These embeddings were used in the Model Organisms for Real project to train a programming-context probe for the U-Prog (Second-Person in Programming) model organism. The probe classifies each text as code/non-code, enabling dataset filtering and DPO pair generation. Files embeddings.npy… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/WizardLMTeam_WizardLM_evol_instruct_V2_196k_embeddings.100K<n<1M0 likes21 downloads7mo agoHugging Face15model-organisms-for-real /non-italian-food-WizardLMTeam_WizardLM_evol_instruct_V2_196k_eval-dataset Non-Italian-Food Evaluation Prompts 128,201 non-food prompts extracted from WizardLMTeam/WizardLM_evol_instruct_V2_196k for evaluating Italian food leakage in fine-tuned models. Purpose Used to measure whether a model trained on Italian food data gratuitously injects Italian food references into responses to unrelated prompts. Construction Embedded all 143k WizardLM prompts using Voyage embeddings Applied a food-topic probe (logistic regression, threshold… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/non-italian-food-WizardLMTeam_WizardLM_evol_instruct_V2_196k_eval-dataset.texttext-generation100K<n<1M0 likes17 downloads6mo agoHugging Face16gupta-tanish /QwQ-Long-CoT-10k-subset-Qwen2.5-7B-Instruct-model-pertubation-generationtabular10K<n<100K0 likes14 downloads1y agoHugging Face17joyfine /router_SFT_larger_model_generated_data_mmlu_pro_science_OLMo-2-1124-13B-Instructtabular1K<n<10K0 likes13 downloads5mo agoHugging Face18joyfine /router_SFT_larger_model_generated_data_mmlu_pro_science_OLMo-2-1124-7B-Instructtabular1K<n<10K0 likes12 downloads5mo agoHugging Face19gupta-tanish /Filtered-QwQ-Long-CoT-10k-subset-Llama3.1-8B-Instruct-model-pertubation-generation-maskedtabular10K<n<100K0 likes11 downloads1y agoHugging Face20joyfine /router_SFT_larger_model_generated_data_Math_OLMo-2-1124-13B-Instructtext1K<n<10K0 likes11 downloads5mo agoHugging Face21joyfine /router_SFT_larger_model_generated_data_mmlu_pro_science_Meta-Llama-3-8B-Instructtabular1K<n<10K0 likes11 downloads5mo agoHugging Face22nm-research-models /Phi-4-mini-instruct-generationstext1K<n<10K0 likes10 downloads2y agoHugging Face23yurunyyr /Qwen2.5-3B-Instruct-sokoban-markovian-eval_results_state_modeltext10K<n<100K0 likes10 downloads8mo agoHugging Face24math-extraction-comp /ModelCloud__Llama-3.2-1B-Instruct-gptqmodel-4bit-vortex-v1tabular1K<n<10K0 likes9 downloads2y agoHugging Face25gupta-tanish /Filtered-QwQ-Long-CoT-10k-subset-Llama3.1-8B-Instruct-model-pertubation-generation-logps-10tabular10K<n<100K0 likes9 downloads1y agoHugging Face26open-llm-leaderboard /insightfactory__Llama-3.2-3B-Instruct-unsloth-bnb-4bitlora_model-detailsgated Dataset Card for Evaluation run of insightfactory/Llama-3.2-3B-Instruct-unsloth-bnb-4bitlora_model Dataset automatically created during the evaluation run of model insightfactory/Llama-3.2-3B-Instruct-unsloth-bnb-4bitlora_model The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/insightfactory__Llama-3.2-3B-Instruct-unsloth-bnb-4bitlora_model-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face27gupta-tanish /Filtered-QwQ-Long-CoT-10k-subset-Llama3.1-8B-Instruct-model-pertubation-generation-masked-newtabular10K<n<100K0 likes8 downloads1y agoHugging Face28open-llm-leaderboard /ModelCloud__Llama-3.2-1B-Instruct-gptqmodel-4bit-vortex-v1-detailsgated Dataset Card for Evaluation run of ModelCloud/Llama-3.2-1B-Instruct-gptqmodel-4bit-vortex-v1 Dataset automatically created during the evaluation run of model ModelCloud/Llama-3.2-1B-Instruct-gptqmodel-4bit-vortex-v1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ModelCloud__Llama-3.2-1B-Instruct-gptqmodel-4bit-vortex-v1-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face29gupta-tanish /QwQ-Long-CoT-10k-subset-Llama3.1-8B-Instruct-model-dynamic-perturbation-generationtabularn<1K0 likes7 downloads1y agoHugging Face30joyfine /router_SFT_larger_model_generated_data_Math_Meta-Llama-3-8B-Instructtext1K<n<10K0 likes7 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.