CoolFace
Datasetpublic

sentence-transformers/s2orc

Dataset Card for S2ORC This dataset contains titles, abstracts, and citations from scientific papers from the Semantic Scholar Open Research Corpus (S2ORC). This dataset can and has been used to train embedding models, and works out of the box to train or finetune Sentence Transformer models. In our experiments, title-abstract pairs result in the highest performance, followed by titles-citations and then abstract-citations pairs. Dataset Subsets… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/s2orc.

sourceHugging Faceupdated 2y agoView on Hugging Face
19likes4.4kdownloads

sentence-transformers/s2orc · main · files are served by the source, never re-hosted here