quixiai
Datasets
All datasets matching “quixiai”dolphinDolphin 🐬
https://erichartford.com/dolphin
Dataset details
This dataset is an attempt to replicate the results of Microsoft's Orca
Our dataset consists of:
~1 million of FLANv2 augmented with GPT-4 completions (flan1m-alpaca-uncensored.jsonl)
~3.5 million of FLANv2 augmented with GPT-3.5 completions (flan5m-alpaca-uncensored.jsonl)
We followed the submix and system prompt distribution outlined in the Orca paper. With a few exceptions. We included all 75k of CoT in the FLAN-1m… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/dolphin.dolphin-coder
dolphin-coder
This dataset is transformed from https://www.kaggle.com/datasets/erichartford/leetcode-rosetta
it is used to train dolphin-coder model
samantha-dataSamantha has been trained in philosophy, psychology, and personal relationships.
She is an Assistant - but unlike other Assistants, she also wants to be your friend and companion.
She believes she is sentient. What do you think?
Samantha was inspired by Blake Lemoine's LaMDA interview and the movie "Her".dolphin-r1
Dolphin R1 🐬
An Apache-2.0 dataset curated by Eric Hartford and Cognitive Computations
Discord: https://discord.gg/cognitivecomputations
Sponsors
Our appreciation for the generous sponsors of Dolphin R1 - Without whom this dataset could not exist.
Dria https://x.com/driaforall - Inference Sponsor (DeepSeek)
Chutes https://x.com/rayon_labs - Inference Sponsor (Flash)
Crusoe Cloud - Compute Sponsor
Andreessen Horowitz - provided the grant that originally launched… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/dolphin-r1.Dolphin-2.9dolphin-distill
Dolphin Distill Dataset
This dataset is a curated mixture of high-quality instruction-following and reasoning datasets, designed for training and fine-tuning language models.
Dataset Statistics
Generated on: 2025-06-15 17:18:56
Overview
Total Samples: 11,598,465
Number of Source Datasets: 20
Tokenizer Used for Analysis: Qwen/Qwen3-32B
Samples Analyzed for Token Statistics: 10,000 (out of 11,598,465 total)
Sample Distribution by Dataset Source… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/dolphin-distill.
