CoolFace
Datasetpublic

slvnwhrl/tenkgnad-clustering-p2p

This dataset can be used as a benchmark for clustering word embeddings for German. The datasets contains news article titles and is based on the dataset of the One Million Posts Corpus and 10kGNAD. It contains 10'275 unique samples, 10 splits with 1'436 to 9'962 samples and 9 unique classes. Splits are built similarly to MTEB's TwentyNewsgroupsClustering. Have a look at German Text Embedding Clustering Benchmark (Github, Paper) for more infos, datasets and evaluation results. If you use this… See the full description on the dataset page: https://huggingface.co/datasets/slvnwhrl/tenkgnad-clustering-p2p.

sourceHugging Facecc-by-nc-sa-4.0updated 2y agoView on Hugging Face
0likes191downloads
settings

This repository belongs to slvnwhrl on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nametenkgnad-clustering-p2p
visibilitypublic
licencecc-by-nc-sa-4.0
gatedno
ownerslvnwhrl
Account settings
slvnwhrl/tenkgnad-clustering-p2p · CoolFace