CoolFace
20 results

slv

SLVMBench /SLVMBench SLVMBench: Skill Learning from Video Memory [NeurIPS 2026 Datasets & Benchmarks Track] Official Hugging Face Dataset repository for SLVMBench, a comprehensive benchmark designed to evaluate whether Video Large Language Models (video-LLMs) can acquire procedural skills from long video memory and apply them to real-time, ongoing tasks under heavy distractor noise. 📂 Repository Structure To optimize download efficiency and bandwidth, this repository contains all… See the full description on the dataset page: https://huggingface.co/datasets/SLVMBench/SLVMBench.videovideo-text-to-text10K<n<100K0 likes732 downloads2mo agoHugging FacezID4si /fineweb-2-slv-edutabular10M<n<100M0 likes655 downloads10mo agoHugging Facetinnel123 /SLV-Set SLV-Set This repository releases the annotation portion of SLV-Set used in the SLVR paper. What is included slv_set: 387,039 region-grounded training examples derived from Visual-CoT. slv_2q: 787,102 two-question training examples where each visual region is paired with two semantically different questions. Data format slv_set Each row contains: dataset: source dataset name. split: split name. question_id: example id. image: relative image path… See the full description on the dataset page: https://huggingface.co/datasets/tinnel123/SLV-Set.textvisual-question-answering1M<n<10M0 likes300 downloads5mo agoHugging Faceyiyic /deu_fin_slv_tur_Latn_traintext1M<n<10M0 likes215 downloads2y agoHugging Faceslvnwhrl /tenkgnad-clustering-p2pThis dataset can be used as a benchmark for clustering word embeddings for German. The datasets contains news article titles and is based on the dataset of the One Million Posts Corpus and 10kGNAD. It contains 10'275 unique samples, 10 splits with 1'436 to 9'962 samples and 9 unique classes. Splits are built similarly to MTEB's TwentyNewsgroupsClustering. Have a look at German Text Embedding Clustering Benchmark (Github, Paper) for more infos, datasets and evaluation results. If you use this… See the full description on the dataset page: https://huggingface.co/datasets/slvnwhrl/tenkgnad-clustering-p2p.textn<1K0 likes169 downloads2y agoHugging Faceslvnwhrl /blurbs-clustering-s2sThis dataset can be used as a benchmark for clustering word embeddings for German. The datasets contains book titles and is based on the dataset from the GermEval 2019 Shared Task on Hierarchical Classification of Blurbs. It contains 17'726 unqiue samples, 28 splits with 177 to 16'425 samples and 4 to 93 unique classes. Splits are built similarly to MTEB's ArxivClusteringS2S. Have a look at German Text Embedding Clustering Benchmark (Github, Paper) for more infos, datasets and evaluation… See the full description on the dataset page: https://huggingface.co/datasets/slvnwhrl/blurbs-clustering-s2s.textn<1K0 likes168 downloads2y agoHugging Face