datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic-medical-conversations-deepseek-v3
🍎 Synthetic Multipersona Doctor Patient Conversations.
Author: Nisten Tahiraj
License: MIT
🧠 Generated by DeepSeek V3 running in full BF16.
🛠️ Done in a way that includes induced errors/obfuscations by the AI patients and friendly rebutals and corrected diagnosis from the AI doctors. This makes the dataset very useful as both training data and retrival systems for reducing hallucinations and increasing the diagnosis quality.
🐧 Conversations… See the full description on the dataset page: https://huggingface.co/datasets/OnDeviceMedNotes/synthetic-medical-conversations-deepseek-v3.healthbench
THE CODE IS CURRENTLY BROKEN BUT THE DATASET IS GOOD!!
HealthBench Implementation for using Opensource Judges
Easy-to-use implementation of OpenAI's HealthBench evaluation benchmark with support for any OpenAI API-compatible model as both the system under test and the judge.
Developed by: Nisten Tahiraj / OnDeviceMednotes
License: MIT
Paper: HealthBench: Evaluating Large Language Models Towards Improved Human Health
Overview
This repository contains tools… See the full description on the dataset page: https://huggingface.co/datasets/OnDeviceMedNotes/healthbench.tulu-v2-sft-mixturehh-rlhf-h4SlimOrcaoasst1_top1_2023-08-25on-device-latency
On-Device Latency Benchmark
Real-world inference latency data for mobile-optimized LLMs, measured on actual phone hardware.
Hardware
Spec
Value
Device
Samsung S20 FE 5G
SoC
Snapdragon 865
RAM
8GB
OS
Android 13
Runtime
llama.cpp (4 threads)
Metrics
tokens_per_sec — Generation speed during inference
latency_ms_per_token — Time per generated token
ram_usage_mb — Peak RAM during inference
file_size_mb — GGUF model file size… See the full description on the dataset page: https://huggingface.co/datasets/dispatchAI/on-device-latency.Federated_Meta-learning_on_wearable_devices
