CoolFace
Datasetpublic

SINAI/ALIA-es-biomedical-synthetic-instructions

Dataset Introduction The ALIA Spanish Biomedical Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in biomedical and healthcare tasks with controlled formats and large-scale supervision. It contains: 639,456 instances 961,073,205 tokens 14 task modalities (clinical diagnosis, patient education, ethical reasoning, document… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical-synthetic-instructions.

sourceHugging Facecc-by-sa-4.0updated 4mo agoView on Hugging Face
0likes93downloads

SINAI/ALIA-es-biomedical-synthetic-instructions · main · files are served by the source, never re-hosted here