CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01QuixiAI /dolphinDolphin 🐬 https://erichartford.com/dolphin Dataset details This dataset is an attempt to replicate the results of Microsoft's Orca Our dataset consists of: ~1 million of FLANv2 augmented with GPT-4 completions (flan1m-alpaca-uncensored.jsonl) ~3.5 million of FLANv2 augmented with GPT-3.5 completions (flan5m-alpaca-uncensored.jsonl) We followed the submix and system prompt distribution outlined in the Orca paper. With a few exceptions. We included all 75k of CoT in the FLAN-1m… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/dolphin.texttext-generation1M<n<10M434 likes1.9k downloads3y agoHugging Face02agentlans /QuixiAI-dolphin-distill Clean QuixiAI/dolphin-distill dataset This is an unofficial, reformatted version of QuixiAI/dolphin-distill. It contains mostly English instruction following and conversation datasets. Major changes: only kept the longest valid conversation from each row (optional system prompt, followed by alternating user and gpt turns) duplicate rows removed URLs, e-mail addresses, phone numbers, API keys and tokens redacted shuffled and split into chunks This filtered the original 11,625,521… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/QuixiAI-dolphin-distill.texttext-generation1M<n<10M0 likes228 downloads10mo agoHugging Face03d0rj /dolphin-ru Dolphin-ru 🐬 This is translated version of ehartford/dolphin into Russian. texttext-classification1M<n<10M9 likes170 downloads3y agoHugging Face04reciperesearch /dolphin-sft-v0.1-preferenceThe preference dataset was generated using Mistral-Instruct-v0.1 finetuned on a GPT-4 subset of the Dolphin dataset (16k samples). Link to the model. Generated responses are labeled as rejected, GPT-4 responses (original Dolphin data) are labeled as accepted. The motivation was to test out the SPIN paper finetuning methodology. texttext-generation10K<n<100K11 likes88 downloads2y agoHugging Face05Maximiliano-Flores-Dev /QuixiAI-dolphin_DatasetDolphin 🐬 https://erichartford.com/dolphin Dataset details This dataset is an attempt to replicate the results of Microsoft's Orca Our dataset consists of: ~1 million of FLANv2 augmented with GPT-4 completions (flan1m-alpaca-uncensored.jsonl) ~3.5 million of FLANv2 augmented with GPT-3.5 completions (flan5m-alpaca-uncensored.jsonl) We followed the submix and system prompt distribution outlined in the Orca paper. With a few exceptions. We included all 75k of CoT in the FLAN-1m… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/QuixiAI-dolphin_Dataset.texttext-generation1M<n<10M1 likes66 downloads3d agoHugging Face06minpeter /dolphin-r1-korean-deepseek-parsed [PARSED] dolphin R1 korean deepseek (toolcalls) The data in this dataset is a subset of the original exp-models/dolphin-r1-korean-deepseek-toolcalls*Dropped row 1273 due to surrogates error. Subset name multi-turn parallel multiple definition Last turn type number of dataset dolphin-r1-korean-deepseek no yes yes tool_calls 1757 dolphin-r1-korean-deepseek-non-reasoning no yes yes tool_calls 1757 This dataset is a re-parsed version of… See the full description on the dataset page: https://huggingface.co/datasets/minpeter/dolphin-r1-korean-deepseek-parsed.texttext-generation1K<n<10K1 likes58 downloads1y agoHugging Face07Hi-Dolphin /MaritimeBench Maritime Bench 本评测集是航运行业首个基于“学科(一级)- 子学科(二级)- 具体考点(三级)”分类体系打造的专业知识评测集,包含1888道客观选择题,覆盖航海、轮机、电子电气员、GMDSS及船员培训等核心领域。评测内容涵盖理论知识、操作技能和行业规范,旨在提升航运领域AI模型的理解与推理能力,确保其在关键知识上的准确性和适应性。同时,本评测集可为航运专业考试、船员培训及资质认证提供自动化测评支持,并优化船舶管理、导航操作、海上通信等场景中的智能问答与决策系统。 MaritimeBench基于行业权威标准,构建了系统、科学的航运知识评测体系,全面评估模型在航海、轮机、电子电气员、GMDSS及船员培训等领域的表现。评测内容深入理论、实践与规范,助力提升AI模型的专业能力。 MaritimeBench评测集亮点 权威性:严格遵循航运行业标准,确保评测科学、实用。 精准分类:采用“学科-子学科-考点”三级框架,评测更具针对性和可扩展性。… See the full description on the dataset page: https://huggingface.co/datasets/Hi-Dolphin/MaritimeBench.texttext-generation1K<n<10K1 likes48 downloads1y agoHugging Face08mayflowergmbh /dolphin_deA german translation for the cognitivecomputations/dolphin dataset. Extracted from seedboxventures/multitask_german_examples_32k. Translation created by seedbox ai for KafkaLM ❤️. Available for finetuning in hiyouga/LLaMA-Factory. texttext-generation10K<n<100K2 likes44 downloads3y agoHugging Face09polymer /dolphin-only-gpt-4Dolphin 🐬 https://erichartford.com/dolphin Dataset details This dataset is an attempt to replicate the results of Microsoft's Orca Our dataset consists of: ~1 million of FLANv2 augmented with GPT-4 completions (flan1m-alpaca-uncensored.jsonl) ~3.5 million of FLANv2 augmented with GPT-3.5 completions (flan5m-alpaca-uncensored.jsonl) We followed the submix and system prompt distribution outlined in the Orca paper. With a few exceptions. We included all 75k of CoT in the FLAN-1m… See the full description on the dataset page: https://huggingface.co/datasets/polymer/dolphin-only-gpt-4.texttext-generation100K<n<1M2 likes38 downloads3y agoHugging Face10mizinovmv /ru_dolphin-r1_v1texttext-generation10K<n<100K0 likes27 downloads1y agoHugging Face11Imunlucky /dolphinDolphin 🐬 https://erichartford.com/dolphin Dataset details This dataset is an attempt to replicate the results of Microsoft's Orca Our dataset consists of: ~1 million of FLANv2 augmented with GPT-4 completions (flan1m-alpaca-uncensored.jsonl) ~3.5 million of FLANv2 augmented with GPT-3.5 completions (flan5m-alpaca-uncensored.jsonl) We followed the submix and system prompt distribution outlined in the Orca paper. With a few exceptions. We included all 75k of CoT in the FLAN-1m… See the full description on the dataset page: https://huggingface.co/datasets/Imunlucky/dolphin.texttext-generation1M<n<10M0 likes20 downloads6mo agoHugging Face12pperojas /dolphinDolphin 🐬 https://erichartford.com/dolphin Dataset details This dataset is an attempt to replicate the results of Microsoft's Orca Our dataset consists of: ~1 million of FLANv2 augmented with GPT-4 completions (flan1m-alpaca-uncensored.jsonl) ~3.5 million of FLANv2 augmented with GPT-3.5 completions (flan5m-alpaca-uncensored.jsonl) We followed the submix and system prompt distribution outlined in the Orca paper. With a few exceptions. We included all 75k of CoT in the FLAN-1m… See the full description on the dataset page: https://huggingface.co/datasets/pperojas/dolphin.texttext-generation1M<n<10M0 likes19 downloads4mo agoHugging Face13tog /dolphin_5k_testTiny Dolphin 🐬 see https://erichartford.com/dolphin Dataset details This dataset is an extract of ~1 million of FLANv2 augmented with GPT-4 completions (flan1m-alpaca-uncensored.jsonl). It is derived from this dataset Loading dataset = load_dataset("tog/dolphin_5k_test) This dataset is licensed apache-2.0 for commercial or non-commercial use. texttext-generation1K<n<10K0 likes18 downloads3y agoHugging Face14erfanzar /dolphin Dataset Card for "Dolphin" This Dataset is edited version of cognitivecomputations/dolphin which only contains conversations and GPT-4 Responses to make it easier to use it with SFTTrinaer texttext-generation100K<n<1M0 likes18 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.