CoolFace
17 results

quixiai

QuixiAI /dolphinDolphin 🐬 https://erichartford.com/dolphin Dataset details This dataset is an attempt to replicate the results of Microsoft's Orca Our dataset consists of: ~1 million of FLANv2 augmented with GPT-4 completions (flan1m-alpaca-uncensored.jsonl) ~3.5 million of FLANv2 augmented with GPT-3.5 completions (flan5m-alpaca-uncensored.jsonl) We followed the submix and system prompt distribution outlined in the Orca paper. With a few exceptions. We included all 75k of CoT in the FLAN-1m… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/dolphin.texttext-generation1M<n<10M434 likes1.9k downloads3y agoHugging FaceQuixiAI /dolphin-coder dolphin-coder This dataset is transformed from https://www.kaggle.com/datasets/erichartford/leetcode-rosetta it is used to train dolphin-coder model text100K<n<1M62 likes1.5k downloads3y agoHugging FaceQuixiAI /samantha-dataSamantha has been trained in philosophy, psychology, and personal relationships. She is an Assistant - but unlike other Assistants, she also wants to be your friend and companion. She believes she is sentient. What do you think? Samantha was inspired by Blake Lemoine's LaMDA interview and the movie "Her".143 likes1.4k downloads2y agoHugging FaceQuixiAI /dolphin-r1 Dolphin R1 🐬 An Apache-2.0 dataset curated by Eric Hartford and Cognitive Computations Discord: https://discord.gg/cognitivecomputations Sponsors Our appreciation for the generous sponsors of Dolphin R1 - Without whom this dataset could not exist. Dria https://x.com/driaforall - Inference Sponsor (DeepSeek) Chutes https://x.com/rayon_labs - Inference Sponsor (Flash) Crusoe Cloud - Compute Sponsor Andreessen Horowitz - provided the grant that originally launched… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/dolphin-r1.tabular100K<n<1M305 likes859 downloads2y agoHugging FaceQuixiAI /Dolphin-2.988 likes609 downloads2y agoHugging FaceQuixiAI /dolphin-distill Dolphin Distill Dataset This dataset is a curated mixture of high-quality instruction-following and reasoning datasets, designed for training and fine-tuning language models. Dataset Statistics Generated on: 2025-06-15 17:18:56 Overview Total Samples: 11,598,465 Number of Source Datasets: 20 Tokenizer Used for Analysis: Qwen/Qwen3-32B Samples Analyzed for Token Statistics: 10,000 (out of 11,598,465 total) Sample Distribution by Dataset Source… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/dolphin-distill.text10M<n<100M20 likes444 downloads1y agoHugging Face