CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlx-community /Apertus-v1.5-QAT-10K mlx-community/Apertus-v1.5-QAT-10K This is a 2000 sample subset of the chosen pairs inside swiss-ai/Apertus_v1p5_Preference_Data for MLX-LM-LoRA and MLX-LoRA-Studio and the Quantization Aware Trained Appertus models. texttext-generation10K<n<100K1 likes120 downloads6d agoHugging Face02ICT-TIME-and-Querit /BOOM-v1.5-training-data Some retrieval datasets of the first stage training are not uploaded: NQ, ELI5, TriviaQA, and MS MARCO document. Please waiting ... Or you can download from the offical website. Citation If you find our work helpful, feel free to give us a cite. @article{zhang2026bagging, title={Bagging-Based Model Merging for Robust General Text Embeddings}, author={Zhang, Hengran and Bi, Keping and Guo, Jiafeng and Zhang, Jiaming and Yang, Wenbo and Shi, Daiting and Cheng… See the full description on the dataset page: https://huggingface.co/datasets/ICT-TIME-and-Querit/BOOM-v1.5-training-data.textsentence-similarity1M<n<10M0 likes76 downloads4mo agoHugging Face03migtissera /Synthia-Coder-v1.5-Itext10K<n<100K34 likes75 downloads2y agoHugging Face04birdsql /usersim-guard-v1.5🌐 Website • 📄 Paper • 💻 BIRD-Interact 🛡️ Overview: USERSIM-GUARD USERSIM-GUARD is a benchmark designed to evaluate the safety, reliability, and robustness of User Simulators in interactive Text-to-SQL environments. As proposed in the BIRD-Interact paper, a high-quality User Simulator must not only be helpful but also "guarded", which means it should provide helpful responses while refusing to leak sensitive information (like solution ideas, database schema details, or… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/usersim-guard-v1.5.text1K<n<10K0 likes51 downloads8mo agoHugging Face05open-llm-leaderboard /Sakalti__ultiima-72B-v1.5-detailsgated Dataset Card for Evaluation run of Sakalti/ultiima-72B-v1.5 Dataset automatically created during the evaluation run of model Sakalti/ultiima-72B-v1.5 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__ultiima-72B-v1.5-details.tabular10K<n<100K0 likes45 downloads2y agoHugging Face06migtissera /Synthia-v1.5-Itext10K<n<100K44 likes43 downloads2y agoHugging Face07migtissera /Synthia-v1.5-IItext10K<n<100K21 likes43 downloads2y agoHugging Face08PCL-Reasoner /V1.5-RL-Math PCL-Reasoner-V1.5 RL Training Dataset Dataset Summary This dataset contains 6,068 unique mathematical reasoning problems extracted from NVIDIA's Nemotron-Post-Training-Dataset-v1. The dataset was specifically curated for reinforcing the mathematical reasoning capabilities of the PCL-Reasoner-V1.5 model through offline reinforcement learning. Each sample includes challenging mathematical problems with long Chain-of-Thought (CoT) reasoning paths exceeding 32K tokens.… See the full description on the dataset page: https://huggingface.co/datasets/PCL-Reasoner/V1.5-RL-Math.text1K<n<10K0 likes40 downloads8mo agoHugging Face09migtissera /Tess-v1.5text100K<n<1M59 likes33 downloads2y agoHugging Face10open-llm-leaderboard /jeffmeloy__Qwen2.5-7B-olm-v1.5-detailsgated Dataset Card for Evaluation run of jeffmeloy/Qwen2.5-7B-olm-v1.5 Dataset automatically created during the evaluation run of model jeffmeloy/Qwen2.5-7B-olm-v1.5 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jeffmeloy__Qwen2.5-7B-olm-v1.5-details.tabular10K<n<100K0 likes32 downloads2y agoHugging Face11florianhoenicke /jina-website-100-64-16-BAAI_bge-small-en-v1.5-1000_9062874564 jina-website-100-64-16-BAAI_bge-small-en-v1.5-1000_9062874564 Dataset Dataset Description jina-website-100-64-16-BAAI_bge-small-en-v1.5-1000_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks. Associated Model This dataset was used to train the jina-website-100-64-16-BAAI_bge-small-en-v1.5-1000_9062874564 model. How to Use To use this dataset for model training or evaluation… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/jina-website-100-64-16-BAAI_bge-small-en-v1.5-1000_9062874564.textn<1K0 likes31 downloads2y agoHugging Face12open-llm-leaderboard /lmsys__vicuna-7b-v1.5-detailsgated Dataset Card for Evaluation run of lmsys/vicuna-7b-v1.5 Dataset automatically created during the evaluation run of model lmsys/vicuna-7b-v1.5 The dataset is composed of 81 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 8 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lmsys__vicuna-7b-v1.5-details.tabular10K<n<100K0 likes30 downloads2y agoHugging Face13lemon-mint /embedded_testdata_nomic_embed_text_v1.5text10K<n<100K0 likes23 downloads2y agoHugging Face14florianhoenicke /jina-website-1-0-16-BAAI_bge-small-en-v1.5-50_9062874564 jina-website-1-0-16-BAAI_bge-small-en-v1.5-50_9062874564 Dataset Dataset Description jina-website-1-0-16-BAAI_bge-small-en-v1.5-50_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks. Associated Model This dataset was used to train the jina-website-1-0-16-BAAI_bge-small-en-v1.5-50_9062874564 model. How to Use To use this dataset for model training or evaluation, you can load it… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/jina-website-1-0-16-BAAI_bge-small-en-v1.5-50_9062874564.textn<1K0 likes20 downloads2y agoHugging Face15florianhoenicke /jina-website-1-0-1-BAAI_bge-small-en-v1.5-50_9062874564 jina-website-1-0-1-BAAI_bge-small-en-v1.5-50_9062874564 Dataset Dataset Description jina-website-1-0-1-BAAI_bge-small-en-v1.5-50_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks. Associated Model This dataset was used to train the jina-website-1-0-1-BAAI_bge-small-en-v1.5-50_9062874564 model. How to Use To use this dataset for model training or evaluation, you can load it… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/jina-website-1-0-1-BAAI_bge-small-en-v1.5-50_9062874564.textn<1K0 likes19 downloads2y agoHugging Face16CreitinGameplays /elisa-chan-v1.5Elisa-chan's dataset generated by ChatGPT "Elisa-chan, an exuberant 20-year-old Japanese woman chatbot! Whether your conversation partner is a fan of games, anime, or just needs a mood lift, you've got the perfect remedy. Encourage them to open up, sharing their thoughts or seeking advice, as you're dedicated to brightening their day. Remind them that if they ever feel a bit low, you're here to effortlessly bring a smile to their face." text1K<n<10K0 likes16 downloads3y agoHugging Face17florianhoenicke /jina-website-1-64-16-BAAI_bge-small-en-v1.5-50_9062874564 jina-website-1-64-16-BAAI_bge-small-en-v1.5-50_9062874564 Dataset Dataset Description jina-website-1-64-16-BAAI_bge-small-en-v1.5-50_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks. Associated Model This dataset was used to train the jina-website-1-64-16-BAAI_bge-small-en-v1.5-50_9062874564 model. How to Use To use this dataset for model training or evaluation, you can load… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/jina-website-1-64-16-BAAI_bge-small-en-v1.5-50_9062874564.textn<1K0 likes14 downloads2y agoHugging Face18juvi21 /Tess-v1.5-sharegpt Tess-v1.5 in sharegpt format Main dataset taken from the great https://huggingface.co/datasets/migtissera/Tess-v1.5 text100K<n<1M0 likes13 downloads2y agoHugging Face19open-llm-leaderboard /jeffmeloy__Qwen2.5-7B-nerd-uncensored-v1.5-detailsgated Dataset Card for Evaluation run of jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.5 Dataset automatically created during the evaluation run of model jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.5 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jeffmeloy__Qwen2.5-7B-nerd-uncensored-v1.5-details.tabular10K<n<100K0 likes13 downloads2y agoHugging Face20open-llm-leaderboard /DoppelReflEx__L3-8B-R1-WolfCore-V1.5-test-detailsgated Dataset Card for Evaluation run of DoppelReflEx/L3-8B-R1-WolfCore-V1.5-test Dataset automatically created during the evaluation run of model DoppelReflEx/L3-8B-R1-WolfCore-V1.5-test The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__L3-8B-R1-WolfCore-V1.5-test-details.tabular10K<n<100K0 likes13 downloads2y agoHugging Face21florianhoenicke /jina-website-100-64-16-BAAI_bge-small-en-v1.5-50_9062874564 jina-website-100-64-16-BAAI_bge-small-en-v1.5-50_9062874564 Dataset Dataset Description jina-website-100-64-16-BAAI_bge-small-en-v1.5-50_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks. Associated Model This dataset was used to train the jina-website-100-64-16-BAAI_bge-small-en-v1.5-50_9062874564 model. How to Use To use this dataset for model training or evaluation, you… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/jina-website-100-64-16-BAAI_bge-small-en-v1.5-50_9062874564.textn<1K0 likes12 downloads2y agoHugging Face22Lux0926 /Deepseek-Coder-7B-Instruct-v1.5-CGPO-10ktabular10K<n<100K0 likes11 downloads11mo agoHugging Face23Denn231 /VV-classifier-2.0-payment-runs-v1.5tabularn<1K0 likes9 downloads21h agoHugging Face24florianhoenicke /pet-shop-1000-64-20-BAAI_bge-small-en-v1.5-1000_9062874564 pet-shop-1000-64-20-BAAI_bge-small-en-v1.5-1000_9062874564 Dataset Dataset Description pet-shop-1000-64-20-BAAI_bge-small-en-v1.5-1000_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks. Associated Model This dataset was used to train the pet-shop-1000-64-20-BAAI_bge-small-en-v1.5-1000_9062874564 model. How to Use To use this dataset for model training or evaluation, you can… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/pet-shop-1000-64-20-BAAI_bge-small-en-v1.5-1000_9062874564.text1K<n<10K0 likes7 downloads2y agoHugging Face25open-llm-leaderboard /sophosympatheia__Midnight-Miqu-70B-v1.5-detailsgated Dataset Card for Evaluation run of sophosympatheia/Midnight-Miqu-70B-v1.5 Dataset automatically created during the evaluation run of model sophosympatheia/Midnight-Miqu-70B-v1.5 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sophosympatheia__Midnight-Miqu-70B-v1.5-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face26open-llm-leaderboard /nothingiisreal__L3.1-8B-Celeste-V1.5-detailsgated Dataset Card for Evaluation run of nothingiisreal/L3.1-8B-Celeste-V1.5 Dataset automatically created during the evaluation run of model nothingiisreal/L3.1-8B-Celeste-V1.5 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/nothingiisreal__L3.1-8B-Celeste-V1.5-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face27open-llm-leaderboard /Quazim0t0__Phi4.Turn.R1Distill_v1.5.1-Tensors-detailsgated Dataset Card for Evaluation run of Quazim0t0/Phi4.Turn.R1Distill_v1.5.1-Tensors Dataset automatically created during the evaluation run of model Quazim0t0/Phi4.Turn.R1Distill_v1.5.1-Tensors The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Quazim0t0__Phi4.Turn.R1Distill_v1.5.1-Tensors-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face28open-llm-leaderboard /NotASI__FineTome-v1.5-Llama3.2-3B-1007-detailsgated Dataset Card for Evaluation run of NotASI/FineTome-v1.5-Llama3.2-3B-1007 Dataset automatically created during the evaluation run of model NotASI/FineTome-v1.5-Llama3.2-3B-1007 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NotASI__FineTome-v1.5-Llama3.2-3B-1007-details.tabular10K<n<100K0 likes6 downloads2y agoHugging Face29open-llm-leaderboard /Pinkstack__PARM-V1.5-base-QwQ-Qwen-2.5-o1-3B-detailsgated Dataset Card for Evaluation run of Pinkstack/PARM-V1.5-base-QwQ-Qwen-2.5-o1-3B Dataset automatically created during the evaluation run of model Pinkstack/PARM-V1.5-base-QwQ-Qwen-2.5-o1-3B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Pinkstack__PARM-V1.5-base-QwQ-Qwen-2.5-o1-3B-details.tabular10K<n<100K0 likes6 downloads2y agoHugging Face30open-llm-leaderboard /NotASI__FineTome-v1.5-Llama3.2-1B-1007-detailsgated Dataset Card for Evaluation run of NotASI/FineTome-v1.5-Llama3.2-1B-1007 Dataset automatically created during the evaluation run of model NotASI/FineTome-v1.5-Llama3.2-1B-1007 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NotASI__FineTome-v1.5-Llama3.2-1B-1007-details.tabular10K<n<100K0 likes5 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.