datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dolphinDolphin 🐬
https://erichartford.com/dolphin
Dataset details
This dataset is an attempt to replicate the results of Microsoft's Orca
Our dataset consists of:
~1 million of FLANv2 augmented with GPT-4 completions (flan1m-alpaca-uncensored.jsonl)
~3.5 million of FLANv2 augmented with GPT-3.5 completions (flan5m-alpaca-uncensored.jsonl)
We followed the submix and system prompt distribution outlined in the Orca paper. With a few exceptions. We included all 75k of CoT in the FLAN-1m… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/dolphin.dolphin-coder
dolphin-coder
This dataset is transformed from https://www.kaggle.com/datasets/erichartford/leetcode-rosetta
it is used to train dolphin-coder model
happy-whale-dolphin-classificationu2-bench
U2-BENCH: Ultrasound Understanding Benchmark
U2-BENCH is the first large-scale benchmark for evaluating Large Vision-Language Models (LVLMs) on ultrasound imaging understanding. It provides a diverse, multi-task dataset curated from 40 licensed sources, covering 15 anatomical regions and 8 clinically inspired tasks across classification, detection, regression, and text generation.
Check the 🌟Leaderboard🌟here: https://dolphin-sound.github.io/u2-bench/… See the full description on the dataset page: https://huggingface.co/datasets/DolphinAI/u2-bench.dolphin-r1
Dolphin R1 🐬
An Apache-2.0 dataset curated by Eric Hartford and Cognitive Computations
Discord: https://discord.gg/cognitivecomputations
Sponsors
Our appreciation for the generous sponsors of Dolphin R1 - Without whom this dataset could not exist.
Dria https://x.com/driaforall - Inference Sponsor (DeepSeek)
Chutes https://x.com/rayon_labs - Inference Sponsor (Flash)
Crusoe Cloud - Compute Sponsor
Andreessen Horowitz - provided the grant that originally launched… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/dolphin-r1.QuixiAI-dolphin-distill
Clean QuixiAI/dolphin-distill dataset
This is an unofficial, reformatted version of QuixiAI/dolphin-distill.
It contains mostly English instruction following and conversation datasets.
Major changes:
only kept the longest valid conversation from each row (optional system prompt, followed by alternating user and gpt turns)
duplicate rows removed
URLs, e-mail addresses, phone numbers, API keys and tokens redacted
shuffled and split into chunks
This filtered the original 11,625,521… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/QuixiAI-dolphin-distill.Dolphin-2.9dolphin-distill
Dolphin Distill Dataset
This dataset is a curated mixture of high-quality instruction-following and reasoning datasets, designed for training and fine-tuning language models.
Dataset Statistics
Generated on: 2025-06-15 17:18:56
Overview
Total Samples: 11,598,465
Number of Source Datasets: 20
Tokenizer Used for Analysis: Qwen/Qwen3-32B
Samples Analyzed for Token Statistics: 10,000 (out of 11,598,465 total)
Sample Distribution by Dataset Source… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/dolphin-distill.details_cognitivecomputations__dolphin-2.6-mistral-7b-dpo-laser
Dataset Card for Evaluation run of cognitivecomputations/dolphin-2.6-mistral-7b-dpo-laser
Dataset automatically created during the evaluation run of model cognitivecomputations/dolphin-2.6-mistral-7b-dpo-laser on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_cognitivecomputations__dolphin-2.6-mistral-7b-dpo-laser.dolphin-ru
Dolphin-ru 🐬
This is translated version of ehartford/dolphin into Russian.
details_cognitivecomputations__dolphin-2.6-mistral-7b
Dataset Card for Evaluation run of cognitivecomputations/dolphin-2.6-mistral-7b
Dataset automatically created during the evaluation run of model cognitivecomputations/dolphin-2.6-mistral-7b on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_cognitivecomputations__dolphin-2.6-mistral-7b.dolphin-2.9.3details_cognitivecomputations__dolphincoder-starcoder2-7bdolphin_r1-testdolphin-combined-distill-coderdetails_ehartford__dolphin-2.1-mistral-7b
Dataset Card for Evaluation run of ehartford/dolphin-2.1-mistral-7b
Dataset Summary
Dataset automatically created during the evaluation run of model ehartford/dolphin-2.1-mistral-7b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ehartford__dolphin-2.1-mistral-7b.dolphindetails_mncai__Llama2-7B-guanaco-dolphin-500
Dataset Card for Evaluation run of mncai/Llama2-7B-guanaco-dolphin-500
Dataset automatically created during the evaluation run of model mncai/Llama2-7B-guanaco-dolphin-500 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_mncai__Llama2-7B-guanaco-dolphin-500.details_HenryJJ__dolphin-2.6-mistral-7b-dpo-orca-v3
Dataset Card for Evaluation run of HenryJJ/dolphin-2.6-mistral-7b-dpo-orca-v3
Dataset automatically created during the evaluation run of model HenryJJ/dolphin-2.6-mistral-7b-dpo-orca-v3 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_HenryJJ__dolphin-2.6-mistral-7b-dpo-orca-v3.details_giannisan__penny5-dolphin-einstein-llama3-dare-ties-chatmldetails_fblgit__UNA-dolphin-2.6-mistral-7b-dpo-laser
Dataset Card for Evaluation run of fblgit/UNA-dolphin-2.6-mistral-7b-dpo-laser
Dataset automatically created during the evaluation run of model fblgit/UNA-dolphin-2.6-mistral-7b-dpo-laser on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_fblgit__UNA-dolphin-2.6-mistral-7b-dpo-laser.dolphin-r1-code-onlydetails_HenryJJ__dolphin-2.6-mistral-7b-dpo-orca-v2
Dataset Card for Evaluation run of HenryJJ/dolphin-2.6-mistral-7b-dpo-orca-v2
Dataset automatically created during the evaluation run of model HenryJJ/dolphin-2.6-mistral-7b-dpo-orca-v2 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_HenryJJ__dolphin-2.6-mistral-7b-dpo-orca-v2.details_ehartford__dolphin-llama-13b
Dataset Card for Evaluation run of ehartford/dolphin-llama-13b
Dataset Summary
Dataset automatically created during the evaluation run of model ehartford/dolphin-llama-13b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ehartford__dolphin-llama-13b.details_HenryJJ__dolphin-2.6-mistral-7b-dpo-orca-v1
Dataset Card for Evaluation run of HenryJJ/dolphin-2.6-mistral-7b-dpo-orca-v1
Dataset automatically created during the evaluation run of model HenryJJ/dolphin-2.6-mistral-7b-dpo-orca-v1 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_HenryJJ__dolphin-2.6-mistral-7b-dpo-orca-v1.details_NotAiLOL__Yi-1.5-dolphin-9B
Dataset Card for Evaluation run of NotAiLOL/Yi-1.5-dolphin-9B
Dataset automatically created during the evaluation run of model NotAiLOL/Yi-1.5-dolphin-9B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_NotAiLOL__Yi-1.5-dolphin-9B.datafdetails_dfurman__llama-2-13b-dolphin-peft
Dataset Card for Evaluation run of dfurman/llama-2-13b-dolphin-peft
Dataset Summary
Dataset automatically created during the evaluation run of model dfurman/llama-2-13b-dolphin-peft on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_dfurman__llama-2-13b-dolphin-peft.details_cognitivecomputations__dolphin-2.9.1-mixtral-1x22bdetails_nasiruddin15__Neural-grok-dolphin-Mistral-7B
