datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MedEmbed-training-triplets-v1
MedEmbed Dataset - v1
Dataset Description
The MedEmbed dataset is a specialized collection of medical and clinical data designed for training and evaluating embedding models in healthcare-related natural language processing (NLP) tasks, particularly information retrieval.
GitHub Repo: https://github.com/abhinand5/MedEmbed
Technical Blog Post: Click here
Dataset Summary
This dataset contains various configurations of medical text data, including corpus text… See the full description on the dataset page: https://huggingface.co/datasets/abhinand/MedEmbed-training-triplets-v1.MedEmbed_COVID_en-vi_tripletsMedEmbed-training-triplets-v1
MedEmbed Dataset - v1
Dataset Description
The MedEmbed dataset is a specialized collection of medical and clinical data designed for training and evaluating embedding models in healthcare-related natural language processing (NLP) tasks, particularly information retrieval.
GitHub Repo: https://github.com/abhinand5/MedEmbed
Technical Blog Post: Click here
Dataset Summary
This dataset contains various configurations of medical text data, including… See the full description on the dataset page: https://huggingface.co/datasets/goofy0422/MedEmbed-training-triplets-v1.
