datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tenkgnad-clustering-p2pThis dataset can be used as a benchmark for clustering word embeddings for German.
The datasets contains news article titles and is based on the dataset of the One Million Posts Corpus and 10kGNAD. It contains 10'275 unique samples, 10 splits with 1'436 to 9'962 samples and 9 unique classes. Splits are built similarly to MTEB's TwentyNewsgroupsClustering.
Have a look at German Text Embedding Clustering Benchmark (Github, Paper) for more infos, datasets and evaluation results.
If you use this… See the full description on the dataset page: https://huggingface.co/datasets/slvnwhrl/tenkgnad-clustering-p2p.blurbs-clustering-s2sThis dataset can be used as a benchmark for clustering word embeddings for German.
The datasets contains book titles and is based on the dataset from the GermEval 2019 Shared Task on Hierarchical Classification of Blurbs. It contains 17'726 unqiue samples, 28 splits with 177 to 16'425 samples and 4 to 93 unique classes. Splits are built similarly to MTEB's ArxivClusteringS2S.
Have a look at German Text Embedding Clustering Benchmark (Github, Paper) for more infos, datasets and evaluation… See the full description on the dataset page: https://huggingface.co/datasets/slvnwhrl/blurbs-clustering-s2s.tenkgnad-clustering-s2sThis dataset can be used as a benchmark for clustering word embeddings for German.
The datasets contains news article titles and is based on the dataset of the One Million Posts Corpus and 10kGNAD. It contains 10'267 unique samples, 10 splits with 1'436 to 9'962 samples and 9 unique classes. Splits are built similarly to MTEB's TwentyNewsgroupsClustering.
Have a look at German Text Embedding Clustering Benchmark (Github, Paper) for more infos, datasets and evaluation results.
If you use this… See the full description on the dataset page: https://huggingface.co/datasets/slvnwhrl/tenkgnad-clustering-s2s.blurbs-clustering-p2pThis dataset can be used as a benchmark for clustering word embeddings for German.
The datasets contains book titles and is based on the dataset from the GermEval 2019 Shared Task on Hierarchical Classification of Blurbs. It contains 18'084 unqiue samples, 28 splits with 177 to 16'425 samples and 4 to 93 unique classes. Splits are built similarly to MTEB's ArxivClusteringP2P.
Have a look at German Text Embedding Clustering Benchmark (Github, Paper) for more infos, datasets and evaluation… See the full description on the dataset page: https://huggingface.co/datasets/slvnwhrl/blurbs-clustering-p2p.slvqa
SLVQA — Streaming Long-Video QA, Perception-Test-anchored
10 × 24-hour streaming videos · 3,568 multiple-choice questions · 3 options each (chance = 33.3%).
SLVQA evaluates whether a model can (1) answer visual questions about a day-long video
that is revealed as a stream (the future is withheld), and (2) do so with O(1) query
latency — the time-to-first-token must not grow with how much video has already streamed.
Unlike LLM-generated video-QA sets, every question here is… See the full description on the dataset page: https://huggingface.co/datasets/treeleaves30760/slvqa.
