datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
healthbench
THE CODE IS CURRENTLY BROKEN BUT THE DATASET IS GOOD!!
HealthBench Implementation for using Opensource Judges
Easy-to-use implementation of OpenAI's HealthBench evaluation benchmark with support for any OpenAI API-compatible model as both the system under test and the judge.
Developed by: Nisten Tahiraj / OnDeviceMednotes
License: MIT
Paper: HealthBench: Evaluating Large Language Models Towards Improved Human Health
Overview
This repository contains tools… See the full description on the dataset page: https://huggingface.co/datasets/OnDeviceMedNotes/healthbench.on-device-latency
On-Device Latency Benchmark
Real-world inference latency data for mobile-optimized LLMs, measured on actual phone hardware.
Hardware
Spec
Value
Device
Samsung S20 FE 5G
SoC
Snapdragon 865
RAM
8GB
OS
Android 13
Runtime
llama.cpp (4 threads)
Metrics
tokens_per_sec — Generation speed during inference
latency_ms_per_token — Time per generated token
ram_usage_mb — Peak RAM during inference
file_size_mb — GGUF model file size… See the full description on the dataset page: https://huggingface.co/datasets/dispatchAI/on-device-latency.
