datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Bones_and_Joints_Benchmark
The Bones and Joints Benchmark
Dataset Overview
We have developed a specialized dataset focused on musculoskeletal disorders, designed to systematically evaluate the clinical capabilities of visual language models (VLMs). The evaluation covers knowledge recall, clinical note interpretation, radiology image interpretation, diagnosis generation and rationale, treatment planning and rationale. The dataset primarily includes multiple-choice questions and open-ended… See the full description on the dataset page: https://huggingface.co/datasets/PUTH2025/Bones_and_Joints_Benchmark.Nemotron-Personas-Vietnam
Nemotron-Personas-Vietnam
Hệ thống AI kết hợp để tạo personas tổng hợp dựa trên phân bố thực tế của Việt Nam
A compound AI approach to personas grounded in real-world distributions
Tổng quan (Overview)
Nemotron-Personas-Vietnam là tập dữ liệu personas được cung cấp dưới dạng mã nguồn mở (CC BY 4.0) dựa trên phân bố nhân khẩu học, địa lý và đặc điểm tính cách của người Việt Nam. Tập dữ liệu phản ánh một cách toàn diện sự phong phú và đặc trưng… See the full description on the dataset page: https://huggingface.co/datasets/bond2bill/Nemotron-Personas-Vietnam.bongo-bernard-de-grunne-tefaf-2011
bongo-bernard-de-grunne-tefaf-2011
Dataset created with PDF2Dataset -- OCR + structure-aware chunking pipeline.
Dataset Summary
Metric
Value
Total chunks
49
Avg chars/chunk
753
Avg images/chunk
1.27
Source files
1
Duplicates removed
0
Quality filtered
0
Schema
Column
Type
Description
chunk_id
string
Unique identifier: filename_chunk_N
text
string
Raw markdown chunk with image refs
text_clean
string
Cleaned… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/bongo-bernard-de-grunne-tefaf-2011.
