CoolFace
Datasetpublic

sentence-transformers/askubuntu-questions

Dataset Card for AskUbuntu Questions The AskUbuntu dataset (Lei et al., 2016) is a collection of preprocessed questions taken from AskUbuntu.com 2014 corpus dump. It also comes with 400*20 mannual annotations, marking pairs of questions as "similar" or "non-similar". The dataset is sourced from the original GitHub repository. This dataset contains all questions from the original source, i.e. the text_tokenized.txt.gz data. See also sentence-transformers/askubuntu for the a… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/askubuntu-questions.

sourceHugging Faceupdated 8mo agoView on Hugging Face
1likes33downloads
4 commits on main
0c9999e8mo ago

Upload dataset

tomaarsen
d20e0df8mo ago

Update README.md

tomaarsen
19d6bf38mo ago

Upload dataset

tomaarsen
50caf038mo ago

initial commit

tomaarsen