CoolFace
Datasetpublic

slvnwhrl/tenkgnad-clustering-s2s

This dataset can be used as a benchmark for clustering word embeddings for German. The datasets contains news article titles and is based on the dataset of the One Million Posts Corpus and 10kGNAD. It contains 10'267 unique samples, 10 splits with 1'436 to 9'962 samples and 9 unique classes. Splits are built similarly to MTEB's TwentyNewsgroupsClustering. Have a look at German Text Embedding Clustering Benchmark (Github, Paper) for more infos, datasets and evaluation results. If you use this… See the full description on the dataset page: https://huggingface.co/datasets/slvnwhrl/tenkgnad-clustering-s2s.

sourceHugging Facecc-by-nc-sa-4.0updated 2y agoView on Hugging Face
0likes196downloads
6 commits on main
1c6223a2y ago

add citation info in readme

slvnwhrl
acaa4ca3y ago

add paper to README

slvnwhrl
6cddbe03y ago

add initial version of readme

slvnwhrl
432066f3y ago

add extraction script

slvnwhrl
be05d3c3y ago

add data

slvnwhrl
cfd04a43y ago

initial commit

slvnwhrl