datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
u2-bench
U2-BENCH: Ultrasound Understanding Benchmark
U2-BENCH is the first large-scale benchmark for evaluating Large Vision-Language Models (LVLMs) on ultrasound imaging understanding. It provides a diverse, multi-task dataset curated from 40 licensed sources, covering 15 anatomical regions and 8 clinically inspired tasks across classification, detection, regression, and text generation.
Check the 🌟Leaderboard🌟here: https://dolphin-sound.github.io/u2-bench/… See the full description on the dataset page: https://huggingface.co/datasets/DolphinAI/u2-bench.dolphin-ru
Dolphin-ru 🐬
This is translated version of ehartford/dolphin into Russian.
MaritimeBench
Maritime Bench
本评测集是航运行业首个基于“学科(一级)- 子学科(二级)- 具体考点(三级)”分类体系打造的专业知识评测集,包含1888道客观选择题,覆盖航海、轮机、电子电气员、GMDSS及船员培训等核心领域。评测内容涵盖理论知识、操作技能和行业规范,旨在提升航运领域AI模型的理解与推理能力,确保其在关键知识上的准确性和适应性。同时,本评测集可为航运专业考试、船员培训及资质认证提供自动化测评支持,并优化船舶管理、导航操作、海上通信等场景中的智能问答与决策系统。
MaritimeBench基于行业权威标准,构建了系统、科学的航运知识评测体系,全面评估模型在航海、轮机、电子电气员、GMDSS及船员培训等领域的表现。评测内容深入理论、实践与规范,助力提升AI模型的专业能力。
MaritimeBench评测集亮点
权威性:严格遵循航运行业标准,确保评测科学、实用。
精准分类:采用“学科-子学科-考点”三级框架,评测更具针对性和可扩展性。… See the full description on the dataset page: https://huggingface.co/datasets/Hi-Dolphin/MaritimeBench.mlabonne_orca-agentinstruct-1M-v1-cleaned-DolphinLabeled
orca-agentinstruct-1M-v1-cleaned DolphinLabeled
Part of the DolphinLabeled series of datasets
Presented by Eric Hartford and Cognitive Computations
The purpose of this dataset is to enable filtering of orca-agentinstruct-1M-v1-cleaned dataset.
The original dataset is mlabonne/orca-agentinstruct-1M-v1-cleaned
(thank you to microsoft and mlabonne)
I have modified the dataset using two scripts.
dedupe.py - removes rows with identical final response.
label.py -… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/mlabonne_orca-agentinstruct-1M-v1-cleaned-DolphinLabeled.Dolphin-3-ShareGPT
orca-agentinstruct-1M-v1-cleaned DolphinLabeled
Part of the DolphinLabeled series of datasets
Presented by Eric Hartford and Cognitive Computations
The purpose of this dataset is to enable filtering of orca-agentinstruct-1M-v1-cleaned dataset.
The original dataset is mlabonne/orca-agentinstruct-1M-v1-cleaned
(thank you to microsoft and mlabonne)
I have modified the dataset using two scripts.
dedupe.py - removes rows with identical final response.
label.py -… See the full description on the dataset page: https://huggingface.co/datasets/SicariusSicariiStuff/Dolphin-3-ShareGPT.Dolphin-DPODolphin
