CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01QuixiAI /dolphinDolphin 🐬 https://erichartford.com/dolphin Dataset details This dataset is an attempt to replicate the results of Microsoft's Orca Our dataset consists of: ~1 million of FLANv2 augmented with GPT-4 completions (flan1m-alpaca-uncensored.jsonl) ~3.5 million of FLANv2 augmented with GPT-3.5 completions (flan5m-alpaca-uncensored.jsonl) We followed the submix and system prompt distribution outlined in the Orca paper. With a few exceptions. We included all 75k of CoT in the FLAN-1m… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/dolphin.texttext-generation1M<n<10M434 likes1.9k downloads3y agoHugging Face02QuixiAI /dolphin-coder dolphin-coder This dataset is transformed from https://www.kaggle.com/datasets/erichartford/leetcode-rosetta it is used to train dolphin-coder model text100K<n<1M62 likes1.5k downloads3y agoHugging Face03QuixiAI /dolphin-r1 Dolphin R1 🐬 An Apache-2.0 dataset curated by Eric Hartford and Cognitive Computations Discord: https://discord.gg/cognitivecomputations Sponsors Our appreciation for the generous sponsors of Dolphin R1 - Without whom this dataset could not exist. Dria https://x.com/driaforall - Inference Sponsor (DeepSeek) Chutes https://x.com/rayon_labs - Inference Sponsor (Flash) Crusoe Cloud - Compute Sponsor Andreessen Horowitz - provided the grant that originally launched… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/dolphin-r1.tabular100K<n<1M305 likes859 downloads2y agoHugging Face04agentlans /QuixiAI-dolphin-distill Clean QuixiAI/dolphin-distill dataset This is an unofficial, reformatted version of QuixiAI/dolphin-distill. It contains mostly English instruction following and conversation datasets. Major changes: only kept the longest valid conversation from each row (optional system prompt, followed by alternating user and gpt turns) duplicate rows removed URLs, e-mail addresses, phone numbers, API keys and tokens redacted shuffled and split into chunks This filtered the original 11,625,521… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/QuixiAI-dolphin-distill.texttext-generation1M<n<10M0 likes228 downloads10mo agoHugging Face05Ewere /dolphin-r1-deepseek-stratifiedThis is a direct copy of mlabonne/dolphin-r1-deepseek with extra stratification information added in post-processing. The strata include: task_type output_length response_style domain complexity There has been no validation done on the strata information, use at your own risk. text100K<n<1M0 likes113 downloads1y agoHugging Face06QuixiAI /OpenCoder-LLM_opc-sft-stage1-DolphinLabeled OpenCoder-LLM SFT DolphinLabeled Part of the DolphinLabeled series of datasets Presented by Eric Hartford and Cognitive Computations The purpose of this dataset is to enable filtering of OpenCoder-LLM SFT dataset. The original dataset is OpenCoder-LLM/opc-sft-stage1 I have modified the dataset using two scripts. dedupe.py - removes rows with identical instruction label.py - adds a "flags" column containing the following boolean values: "refusal": whether the… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/OpenCoder-LLM_opc-sft-stage1-DolphinLabeled.text1M<n<10M12 likes92 downloads2y agoHugging Face07reciperesearch /dolphin-sft-v0.1-preferenceThe preference dataset was generated using Mistral-Instruct-v0.1 finetuned on a GPT-4 subset of the Dolphin dataset (16k samples). Link to the model. Generated responses are labeled as rejected, GPT-4 responses (original Dolphin data) are labeled as accepted. The motivation was to test out the SPIN paper finetuning methodology. texttext-generation10K<n<100K11 likes88 downloads2y agoHugging Face08nyu-dice-lab /lm-eval-results-cognitivecomputations-dolphin-2.9-llama3-8b-private Dataset Card for Evaluation run of cognitivecomputations/dolphin-2.9-llama3-8b Dataset automatically created during the evaluation run of model cognitivecomputations/dolphin-2.9-llama3-8b The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-cognitivecomputations-dolphin-2.9-llama3-8b-private.tabular100K<n<1M0 likes77 downloads2y agoHugging Face09Maximiliano-Flores-Dev /QuixiAI-dolphin_DatasetDolphin 🐬 https://erichartford.com/dolphin Dataset details This dataset is an attempt to replicate the results of Microsoft's Orca Our dataset consists of: ~1 million of FLANv2 augmented with GPT-4 completions (flan1m-alpaca-uncensored.jsonl) ~3.5 million of FLANv2 augmented with GPT-3.5 completions (flan5m-alpaca-uncensored.jsonl) We followed the submix and system prompt distribution outlined in the Orca paper. With a few exceptions. We included all 75k of CoT in the FLAN-1m… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/QuixiAI-dolphin_Dataset.texttext-generation1M<n<10M1 likes66 downloads3d agoHugging Face10QuixiAI /OpenCoder-LLM_opc-sft-stage2-DolphinLabeled OpenCoder-LLM SFT DolphinLabeled Part of the DolphinLabeled series of datasets Presented by Eric Hartford and Cognitive Computations The purpose of this dataset is to enable filtering of OpenCoder-LLM SFT dataset. The original dataset is OpenCoder-LLM/opc-sft-stage2 I have modified the dataset using two scripts. dedupe.py - removes rows with identical instruction label.py - adds a "flags" column containing the following boolean values: "refusal": whether the… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/OpenCoder-LLM_opc-sft-stage2-DolphinLabeled.text100K<n<1M8 likes64 downloads2y agoHugging Face11Skorcht /dolphin2.9text100K<n<1M0 likes62 downloads2y agoHugging Face12joeyzero /dolphin-r1-backfill-0.0.2text100K<n<1M0 likes52 downloads11mo agoHugging Face13QuixiAI /allenai_tulu-3-sft-mixture-DolphinLabeled allenai tulu-3-sft-mixture DolphinLabeled Part of the DolphinLabeled series of datasets Presented by Eric Hartford and Cognitive Computations The purpose of this dataset is to enable filtering of allenai/tulu-3-sft-mixture dataset. The original dataset is allenai/tulu-3-sft-mixture I have modified the dataset using two scripts. dedupe.py - removes rows with identical final message content label.py - adds a "flags" column containing the following boolean… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/allenai_tulu-3-sft-mixture-DolphinLabeled.textother100K<n<1M8 likes51 downloads2y agoHugging Face14Hi-Dolphin /MaritimeBench Maritime Bench 本评测集是航运行业首个基于“学科(一级)- 子学科(二级)- 具体考点(三级)”分类体系打造的专业知识评测集,包含1888道客观选择题,覆盖航海、轮机、电子电气员、GMDSS及船员培训等核心领域。评测内容涵盖理论知识、操作技能和行业规范,旨在提升航运领域AI模型的理解与推理能力,确保其在关键知识上的准确性和适应性。同时,本评测集可为航运专业考试、船员培训及资质认证提供自动化测评支持,并优化船舶管理、导航操作、海上通信等场景中的智能问答与决策系统。 MaritimeBench基于行业权威标准,构建了系统、科学的航运知识评测体系,全面评估模型在航海、轮机、电子电气员、GMDSS及船员培训等领域的表现。评测内容深入理论、实践与规范,助力提升AI模型的专业能力。 MaritimeBench评测集亮点 权威性:严格遵循航运行业标准,确保评测科学、实用。 精准分类:采用“学科-子学科-考点”三级框架,评测更具针对性和可扩展性。… See the full description on the dataset page: https://huggingface.co/datasets/Hi-Dolphin/MaritimeBench.texttext-generation1K<n<10K1 likes48 downloads1y agoHugging Face15mayflowergmbh /dolphin_deA german translation for the cognitivecomputations/dolphin dataset. Extracted from seedboxventures/multitask_german_examples_32k. Translation created by seedbox ai for KafkaLM ❤️. Available for finetuning in hiyouga/LLaMA-Factory. texttext-generation10K<n<100K2 likes44 downloads3y agoHugging Face16polymer /dolphin-only-gpt-4Dolphin 🐬 https://erichartford.com/dolphin Dataset details This dataset is an attempt to replicate the results of Microsoft's Orca Our dataset consists of: ~1 million of FLANv2 augmented with GPT-4 completions (flan1m-alpaca-uncensored.jsonl) ~3.5 million of FLANv2 augmented with GPT-3.5 completions (flan5m-alpaca-uncensored.jsonl) We followed the submix and system prompt distribution outlined in the Orca paper. With a few exceptions. We included all 75k of CoT in the FLAN-1m… See the full description on the dataset page: https://huggingface.co/datasets/polymer/dolphin-only-gpt-4.texttext-generation100K<n<1M2 likes38 downloads3y agoHugging Face17open-llm-leaderboard /cognitivecomputations__Dolphin3.0-Llama3.2-1B-detailsgated Dataset Card for Evaluation run of cognitivecomputations/Dolphin3.0-Llama3.2-1B Dataset automatically created during the evaluation run of model cognitivecomputations/Dolphin3.0-Llama3.2-1B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cognitivecomputations__Dolphin3.0-Llama3.2-1B-details.tabular10K<n<100K0 likes38 downloads2y agoHugging Face18open-llm-leaderboard /cognitivecomputations__dolphin-2.9.1-yi-1.5-9b-detailsgated Dataset Card for Evaluation run of cognitivecomputations/dolphin-2.9.1-yi-1.5-9b Dataset automatically created during the evaluation run of model cognitivecomputations/dolphin-2.9.1-yi-1.5-9b The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cognitivecomputations__dolphin-2.9.1-yi-1.5-9b-details.tabular10K<n<100K0 likes37 downloads2y agoHugging Face19QuixiAI /HuggingFaceTB_smoltalk-DolphinLabeled HuggingFaceTB smoltalk DolphinLabeled Part of the DolphinLabeled series of datasets Presented by Eric Hartford and Cognitive Computations The purpose of this dataset is to enable filtering of HuggingFaceTB/smoltalk dataset. The original dataset is HuggingFaceTB/smoltalk I have modified the dataset using two scripts. dedupe.py - removes rows with identical final message content label.py - adds a "flags" column containing the following boolean values: "refusal":… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/HuggingFaceTB_smoltalk-DolphinLabeled.text1M<n<10M10 likes37 downloads2y agoHugging Face20open-llm-leaderboard /cognitivecomputations__dolphin-2.9.3-mistral-7B-32k-detailsgated Dataset Card for Evaluation run of cognitivecomputations/dolphin-2.9.3-mistral-7B-32k Dataset automatically created during the evaluation run of model cognitivecomputations/dolphin-2.9.3-mistral-7B-32k The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cognitivecomputations__dolphin-2.9.3-mistral-7B-32k-details.tabular10K<n<100K0 likes34 downloads2y agoHugging Face21open-llm-leaderboard /NikolaSigmoid__AceMath-1.5B-Instruct-dolphin-r1-200-detailsgated Dataset Card for Evaluation run of NikolaSigmoid/AceMath-1.5B-Instruct-dolphin-r1-200 Dataset automatically created during the evaluation run of model NikolaSigmoid/AceMath-1.5B-Instruct-dolphin-r1-200 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NikolaSigmoid__AceMath-1.5B-Instruct-dolphin-r1-200-details.tabular10K<n<100K0 likes32 downloads2y agoHugging Face22open-llm-leaderboard /cognitivecomputations__dolphin-2.9.3-Yi-1.5-34B-32k-detailsgated Dataset Card for Evaluation run of cognitivecomputations/dolphin-2.9.3-Yi-1.5-34B-32k Dataset automatically created during the evaluation run of model cognitivecomputations/dolphin-2.9.3-Yi-1.5-34B-32k The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cognitivecomputations__dolphin-2.9.3-Yi-1.5-34B-32k-details.tabular10K<n<100K0 likes31 downloads2y agoHugging Face23open-llm-leaderboard /cognitivecomputations__dolphin-2.9.1-llama-3-70b-detailsgated Dataset Card for Evaluation run of cognitivecomputations/dolphin-2.9.1-llama-3-70b Dataset automatically created during the evaluation run of model cognitivecomputations/dolphin-2.9.1-llama-3-70b The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cognitivecomputations__dolphin-2.9.1-llama-3-70b-details.tabular10K<n<100K0 likes31 downloads2y agoHugging Face24open-llm-leaderboard /cognitivecomputations__Dolphin3.0-Qwen2.5-0.5B-detailsgated Dataset Card for Evaluation run of cognitivecomputations/Dolphin3.0-Qwen2.5-0.5B Dataset automatically created during the evaluation run of model cognitivecomputations/Dolphin3.0-Qwen2.5-0.5B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cognitivecomputations__Dolphin3.0-Qwen2.5-0.5B-details.tabular10K<n<100K0 likes31 downloads2y agoHugging Face25open-llm-leaderboard /cognitivecomputations__dolphin-2.9.2-qwen2-72b-detailsgated Dataset Card for Evaluation run of cognitivecomputations/dolphin-2.9.2-qwen2-72b Dataset automatically created during the evaluation run of model cognitivecomputations/dolphin-2.9.2-qwen2-72b The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cognitivecomputations__dolphin-2.9.2-qwen2-72b-details.tabular10K<n<100K0 likes29 downloads2y agoHugging Face26PJMixers-Dev /cognitivecomputations_dolphin-r1-reasoning-deepseektext100K<n<1M0 likes29 downloads2y agoHugging Face27QuixiAI /mlabonne_orca-agentinstruct-1M-v1-cleaned-DolphinLabeled orca-agentinstruct-1M-v1-cleaned DolphinLabeled Part of the DolphinLabeled series of datasets Presented by Eric Hartford and Cognitive Computations The purpose of this dataset is to enable filtering of orca-agentinstruct-1M-v1-cleaned dataset. The original dataset is mlabonne/orca-agentinstruct-1M-v1-cleaned (thank you to microsoft and mlabonne) I have modified the dataset using two scripts. dedupe.py - removes rows with identical final response. label.py -… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/mlabonne_orca-agentinstruct-1M-v1-cleaned-DolphinLabeled.textquestion-answering1M<n<10M6 likes28 downloads2y agoHugging Face28open-llm-leaderboard /cognitivecomputations__dolphin-2.9-llama3-8b-detailsgated Dataset Card for Evaluation run of cognitivecomputations/dolphin-2.9-llama3-8b Dataset automatically created during the evaluation run of model cognitivecomputations/dolphin-2.9-llama3-8b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cognitivecomputations__dolphin-2.9-llama3-8b-details.tabular10K<n<100K0 likes27 downloads2y agoHugging Face29open-llm-leaderboard /cognitivecomputations__dolphin-2.9.2-Phi-3-Medium-abliterated-detailsgated Dataset Card for Evaluation run of cognitivecomputations/dolphin-2.9.2-Phi-3-Medium-abliterated Dataset automatically created during the evaluation run of model cognitivecomputations/dolphin-2.9.2-Phi-3-Medium-abliterated The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cognitivecomputations__dolphin-2.9.2-Phi-3-Medium-abliterated-details.tabular10K<n<100K1 likes26 downloads2y agoHugging Face30open-llm-leaderboard /cognitivecomputations__dolphin-2.9.2-qwen2-7b-detailsgated Dataset Card for Evaluation run of cognitivecomputations/dolphin-2.9.2-qwen2-7b Dataset automatically created during the evaluation run of model cognitivecomputations/dolphin-2.9.2-qwen2-7b The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cognitivecomputations__dolphin-2.9.2-qwen2-7b-details.tabular10K<n<100K0 likes25 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.