CoolFace
Datasetpublic

sentence-transformers/askubuntu

Dataset Card for AskUbuntu The AskUbuntu dataset (Lei et al., 2016) is a collection of preprocessed questions taken from AskUbuntu.com 2014 corpus dump. It also comes with 400*20 mannual annotations, marking pairs of questions as "similar" or "non-similar". The dataset is sourced from the original GitHub repository. Note that for the train split, the "positive" is the list of similar questions according to AskUbuntu, and "negative" is a list of randomly selected questions. For… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/askubuntu.

sourceHugging Faceupdated 8mo agoView on Hugging Face
2likes28downloads
../
filedev-00000-of-00001.parquet129 KBdownload
filetest-00000-of-00001.parquet131 KBdownload
filetrain-00000-of-00001.parquet40.5 MBdownload

sentence-transformers/askubuntu · main · files are served by the source, never re-hosted here