CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bio-nlp-umass /bioinstruct Dataset Card for BioInstruct GitHub repo: https://github.com/bio-nlp/BioInstruct Dataset Summary BioInstruct is a dataset of 25k instructions and demonstrations generated by OpenAI's GPT-4 engine in July 2023. This instruction data can be used to conduct instruction-tuning for language models (e.g. Llama) and make the language model follow biomedical instruction better. Improvements of Llama on 9 common BioMedical tasks are shown in the result section. Taking… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/bioinstruct.texttext-generation10K<n<100K25 likes115 downloads2y agoHugging Face02bio-nlp-umass /MedQA-MM MedQA-MM Identifier Release Paper repository · Hugging Face dataset MedQA-MM is a 1,000-item shortcut-mitigated medical multimodal multiple-choice benchmark constructed from MedThinkVQA, MedXpertQA-MM, and the Health and Medicine portion of MMMU. This public release is intentionally identifier-only. It does not contain source questions, answer choices, gold answers, images, clinical text, or repaired payloads. It provides stable source locators, pinned source revisions, and a… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/MedQA-MM.tabularvisual-question-answering1K<n<10K0 likes102 downloads21d agoHugging Face03bio-nlp-umass /MedQA-CS-ExamBenchmarking LLMs Clinical Skills for Patient-Centered Diagnostics and Documentation Project github: https://github.com/bio-nlp/MedQA-CS MedQA-CS-Student dataset: https://huggingface.co/datasets/bio-nlp-umass/MedQA-CS-Student tabularquestion-answering1K<n<10K10 likes42 downloads2y agoHugging Face04bio-nlp-umass /MedQA-CS-Studenttabular1K<n<10K3 likes21 downloads2y agoHugging Face05automated-analytics /bionlp2004text10K<n<100K0 likes6 downloads1y agoHugging Face06DUTIR-BioNLP /Taiyi_Instruction_Data_001 Taiyi_Instruction_Data_001 The raw instruction data is used to train Taiyi LLM. The data is distributed under CC BY-NC-SA 4.0. The original benchmark datasets that support this study are available from the official websites of natural language processing challenges with Data Use Agreements. More details can be found in Taiyi project. Citation If you use the repository of this project, please cite it. @article{Taiyi, title="{Taiyi: A Bilingual Fine-Tuned Large… See the full description on the dataset page: https://huggingface.co/datasets/DUTIR-BioNLP/Taiyi_Instruction_Data_001.text1M<n<10M0 likes5 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.