CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SLVMBench /SLVMBench SLVMBench: Skill Learning from Video Memory [NeurIPS 2026 Datasets & Benchmarks Track] Official Hugging Face Dataset repository for SLVMBench, a comprehensive benchmark designed to evaluate whether Video Large Language Models (video-LLMs) can acquire procedural skills from long video memory and apply them to real-time, ongoing tasks under heavy distractor noise. 📂 Repository Structure To optimize download efficiency and bandwidth, this repository contains all… See the full description on the dataset page: https://huggingface.co/datasets/SLVMBench/SLVMBench.videovideo-text-to-text10K<n<100K0 likes732 downloads2mo agoHugging Face02zID4si /fineweb-2-slv-edutabular10M<n<100M0 likes655 downloads10mo agoHugging Face03tinnel123 /SLV-Set SLV-Set This repository releases the annotation portion of SLV-Set used in the SLVR paper. What is included slv_set: 387,039 region-grounded training examples derived from Visual-CoT. slv_2q: 787,102 two-question training examples where each visual region is paired with two semantically different questions. Data format slv_set Each row contains: dataset: source dataset name. split: split name. question_id: example id. image: relative image path… See the full description on the dataset page: https://huggingface.co/datasets/tinnel123/SLV-Set.textvisual-question-answering1M<n<10M0 likes300 downloads5mo agoHugging Face04yiyic /deu_fin_slv_tur_Latn_traintext1M<n<10M0 likes215 downloads2y agoHugging Face05slvnwhrl /tenkgnad-clustering-p2pThis dataset can be used as a benchmark for clustering word embeddings for German. The datasets contains news article titles and is based on the dataset of the One Million Posts Corpus and 10kGNAD. It contains 10'275 unique samples, 10 splits with 1'436 to 9'962 samples and 9 unique classes. Splits are built similarly to MTEB's TwentyNewsgroupsClustering. Have a look at German Text Embedding Clustering Benchmark (Github, Paper) for more infos, datasets and evaluation results. If you use this… See the full description on the dataset page: https://huggingface.co/datasets/slvnwhrl/tenkgnad-clustering-p2p.textn<1K0 likes169 downloads2y agoHugging Face06slvnwhrl /blurbs-clustering-s2sThis dataset can be used as a benchmark for clustering word embeddings for German. The datasets contains book titles and is based on the dataset from the GermEval 2019 Shared Task on Hierarchical Classification of Blurbs. It contains 17'726 unqiue samples, 28 splits with 177 to 16'425 samples and 4 to 93 unique classes. Splits are built similarly to MTEB's ArxivClusteringS2S. Have a look at German Text Embedding Clustering Benchmark (Github, Paper) for more infos, datasets and evaluation… See the full description on the dataset page: https://huggingface.co/datasets/slvnwhrl/blurbs-clustering-s2s.textn<1K0 likes168 downloads2y agoHugging Face07slvnwhrl /tenkgnad-clustering-s2sThis dataset can be used as a benchmark for clustering word embeddings for German. The datasets contains news article titles and is based on the dataset of the One Million Posts Corpus and 10kGNAD. It contains 10'267 unique samples, 10 splits with 1'436 to 9'962 samples and 9 unique classes. Splits are built similarly to MTEB's TwentyNewsgroupsClustering. Have a look at German Text Embedding Clustering Benchmark (Github, Paper) for more infos, datasets and evaluation results. If you use this… See the full description on the dataset page: https://huggingface.co/datasets/slvnwhrl/tenkgnad-clustering-s2s.textn<1K0 likes166 downloads2y agoHugging Face08slvnwhrl /blurbs-clustering-p2pThis dataset can be used as a benchmark for clustering word embeddings for German. The datasets contains book titles and is based on the dataset from the GermEval 2019 Shared Task on Hierarchical Classification of Blurbs. It contains 18'084 unqiue samples, 28 splits with 177 to 16'425 samples and 4 to 93 unique classes. Splits are built similarly to MTEB's ArxivClusteringP2P. Have a look at German Text Embedding Clustering Benchmark (Github, Paper) for more infos, datasets and evaluation… See the full description on the dataset page: https://huggingface.co/datasets/slvnwhrl/blurbs-clustering-p2p.textn<1K0 likes163 downloads2y agoHugging Face09slvrfivo /physicalaiThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/slvrfivo/physicalai.tabularrobotics1K<n<10K0 likes101 downloads11d agoHugging Face10tohoku-nlp /SLVMEval Dataset Card for SLVMEval Dataset Summary SLVMEval (Synthetic Long-Video Meta-Evaluation Benchmark) is a benchmark for meta-evaluating automatic evaluation systems for text-to-long video (T2LV) generation. The benchmark follows a pairwise comparison-based setup. It constructs controlled high-quality vs. low-quality long-video pairs by applying aspect-specific synthetic degradations to source videos. The final benchmark data is built by retaining human-validated pairs… See the full description on the dataset page: https://huggingface.co/datasets/tohoku-nlp/SLVMEval.image1K<n<10K2 likes41 downloads7mo agoHugging Face11EdonFetaji /slvesnik-mk-sq Службен весник MK–SQ Legal Parallel Corpus 293,612 Macedonian–Albanian sentence pairs from the Official Gazette of the Republic of North Macedonia (Службен весник), 2001–2025. from datasets import load_dataset ds = load_dataset("EdonFetaji/slvesnik-mk-sq") print(ds["test"][0]["mk_text"], ds["test"][0]["sq_text"]) Splits The split is issue-disjoint: assignment happens at the level of the source PDF, so no gazette issue contributes to more than one split. Gazette… See the full description on the dataset page: https://huggingface.co/datasets/EdonFetaji/slvesnik-mk-sq.tabulartranslation100K<n<1M0 likes34 downloads5d agoHugging Face12AS1ingin /slvhrrimagen<1K0 likes29 downloads1y agoHugging Face13hafeezjimoh /eval_slvla_pick_cubesThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 1, "total_frames": 872, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hafeezjimoh/eval_slvla_pick_cubes.tabularroboticsn<1K0 likes28 downloads8mo agoHugging Face14yiyic /oscar_slv_Latn_traintext10K<n<100K0 likes22 downloads2y agoHugging Face15slvss /spam.csv0 likes20 downloads1y agoHugging Face16yiyic /oscar_slv_Latn_devtextn<1K0 likes17 downloads2y agoHugging Face17zhangyuze999 /eval_slvaThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so_follower", "total_episodes": 0, "total_frames": 0, "total_tasks": 0, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": {}, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zhangyuze999/eval_slva.robotics0 likes12 downloads4mo agoHugging Face18yiyic /oscar_slv_Latn_testtextn<1K0 likes10 downloads2y agoHugging Face19treeleaves30760 /slvqagated SLVQA — Streaming Long-Video QA, Perception-Test-anchored 10 × 24-hour streaming videos · 3,568 multiple-choice questions · 3 options each (chance = 33.3%). SLVQA evaluates whether a model can (1) answer visual questions about a day-long video that is revealed as a stream (the future is withheld), and (2) do so with O(1) query latency — the time-to-first-token must not grow with how much video has already streamed. Unlike LLM-generated video-QA sets, every question here is… See the full description on the dataset page: https://huggingface.co/datasets/treeleaves30760/slvqa.tabularvisual-question-answering1K<n<10K0 likes10 downloads1mo agoHugging Face20SLV-Polynome /slv_videosvideon<1K0 likes7 downloads2y agoHugging Face21slvss /uas0 likes2 downloads1y agoHugging Face22branague01 /SLvsi8EJ0 likes1 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.