CoolFace
Datasetpublic

nuvocare/WikiMedical_sentence_similarity

Dataset Card for "WikiMedical_sentence_similarity" WikiMedical_sentence_similarity is an adapted and ready-to-use sentence similarity dataset based on this dataset. The preprocessing followed three steps: Each text is splitted into sentences of 256 tokens (nltk tokenizer) Each sentence is paired with a positive pair if found, and a negative one. Negative one are drawn randomly in the whole dataset. Train and test split correspond to 70%/30% More Information needed

sourceHugging Faceupdated 3y agoView on Hugging Face
1likes94downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
nuvocare/WikiMedical_sentence_similarity · CoolFace