dolphin
Datasets
All datasets matching “dolphin”dolphinDolphin 🐬
https://erichartford.com/dolphin
Dataset details
This dataset is an attempt to replicate the results of Microsoft's Orca
Our dataset consists of:
~1 million of FLANv2 augmented with GPT-4 completions (flan1m-alpaca-uncensored.jsonl)
~3.5 million of FLANv2 augmented with GPT-3.5 completions (flan5m-alpaca-uncensored.jsonl)
We followed the submix and system prompt distribution outlined in the Orca paper. With a few exceptions. We included all 75k of CoT in the FLAN-1m… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/dolphin.dolphin-coder
dolphin-coder
This dataset is transformed from https://www.kaggle.com/datasets/erichartford/leetcode-rosetta
it is used to train dolphin-coder model
happy-whale-dolphin-classificationu2-bench
U2-BENCH: Ultrasound Understanding Benchmark
U2-BENCH is the first large-scale benchmark for evaluating Large Vision-Language Models (LVLMs) on ultrasound imaging understanding. It provides a diverse, multi-task dataset curated from 40 licensed sources, covering 15 anatomical regions and 8 clinically inspired tasks across classification, detection, regression, and text generation.
Check the 🌟Leaderboard🌟here: https://dolphin-sound.github.io/u2-bench/… See the full description on the dataset page: https://huggingface.co/datasets/DolphinAI/u2-bench.dolphin-r1
Dolphin R1 🐬
An Apache-2.0 dataset curated by Eric Hartford and Cognitive Computations
Discord: https://discord.gg/cognitivecomputations
Sponsors
Our appreciation for the generous sponsors of Dolphin R1 - Without whom this dataset could not exist.
Dria https://x.com/driaforall - Inference Sponsor (DeepSeek)
Chutes https://x.com/rayon_labs - Inference Sponsor (Flash)
Crusoe Cloud - Compute Sponsor
Andreessen Horowitz - provided the grant that originally launched… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/dolphin-r1.QuixiAI-dolphin-distill
Clean QuixiAI/dolphin-distill dataset
This is an unofficial, reformatted version of QuixiAI/dolphin-distill.
It contains mostly English instruction following and conversation datasets.
Major changes:
only kept the longest valid conversation from each row (optional system prompt, followed by alternating user and gpt turns)
duplicate rows removed
URLs, e-mail addresses, phone numbers, API keys and tokens redacted
shuffled and split into chunks
This filtered the original 11,625,521… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/QuixiAI-dolphin-distill.
