CoolFace
Datasetpublic

sentence-transformers/s2orc

Dataset Card for S2ORC This dataset contains titles, abstracts, and citations from scientific papers from the Semantic Scholar Open Research Corpus (S2ORC). This dataset can and has been used to train embedding models, and works out of the box to train or finetune Sentence Transformer models. In our experiments, title-abstract pairs result in the highest performance, followed by titles-citations and then abstract-citations pairs. Dataset Subsets… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/s2orc.

sourceHugging Faceupdated 2y agoView on Hugging Face
19likes4.4kdownloads
settings

This repository belongs to sentence-transformers on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

names2orc
visibilitypublic
licencenot set
gatedno
ownersentence-transformers
Account settings
sentence-transformers/s2orc · CoolFace