datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
omni-med-vqa-miniomni-med-vqa-mini-robustnessvqa-medvqa-med-robustnessomni-med-vqa-mini-v2vqa-med-2019omni-med-vqamed_paragraph_simplificationMedSimp-JudgeBench
MedSimp-JudgeBench
A perturbation-based calibration benchmark for LLM-as-judge medical faithfulness evaluation.
Description
708 reference simplifications with programmatically injected known errors, used to measure
sensitivity and specificity of dual LLM judges (Llama-3.3-70B and Qwen3-32B) via Nebius Token Factory.
Error Types
Error type
n
Llama sensitivity
Qwen sensitivity
Dose 10x
95
0.44
0.80
Lateral swap
150
0.43
0.83
Negation… See the full description on the dataset page: https://huggingface.co/datasets/chambul/MedSimp-JudgeBench.MedsiML_EN_HIN_Data
MedSiML
Dataset Description
This dataset contains simplified English and simplified Hindi sentence pairs derived from the MedSiML dataset. The data has been filtered using the Cynical Data Selection algorithm to retain examples that are most representative of the target distribution while reducing redundancy.
Source Dataset
This dataset is derived from the MedSiML dataset:
MedSiML: A Multilingual Approach for Simplifying Medical Texts… See the full description on the dataset page: https://huggingface.co/datasets/vishnu-vizz/MedsiML_EN_HIN_Data.med-orders-simord-v1meds-imagesii_med_ckeck_simpleqamed-orders-simord-v2
