CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dascim /GreekMMLU GreekMMLU GreekMMLU is a native-sourced benchmark for evaluating massive multitask language understanding in Greek, built from authentic Greek exam-style multiple-choice questions (MCQ) rather than machine-translated English benchmarks. 21,805 questions across 45 subjects 4 high-level groups: STEM, Humanities, Social Sciences, Other Difficulty/education levels spanning Primary → Secondary → University → Professional (+ an N/A bucket) Public vs. private split for… See the full description on the dataset page: https://huggingface.co/datasets/dascim/GreekMMLU.textquestion-answering10K<n<100K7 likes3.4k downloads8mo agoHugging Face02quinnlue /audioset-dasheng-0.6b-emb AudioSet DaSheng-0.6B embeddings Mean-pooled, float16 embeddings of danjacobellis/audioset_opus_24kbps from mispeech/dasheng-0.6B. Columns path: source clip path (string) label: source AudioSet label indices (list of int64) emb: 1,280-dimensional fixed-size list of float16 Audio is decoded from the source Opus bytes, mixed to mono, and resampled to 16 kHz. The embedding is the model's documented outputdim=None output: sigmoid applied to the mean of the final… See the full description on the dataset page: https://huggingface.co/datasets/quinnlue/audioset-dasheng-0.6b-emb.textfeature-extraction1M<n<10M0 likes234 downloads2mo agoHugging Face03dasyd /quants QuAnTS: Question Answering on Time Series QuAnTS is a challenging dataset designed to bridge the gap in question-answering research on time series data. The dataset features a wide variety of questions and answers concerning human movements, presented as tracked skeleton trajectories. QuAnTS also includes human reference performance to benchmark the practical usability of models trained on this dataset. At present, there is no official leaderboard for this dataset. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/dasyd/quants.tabularquestion-answering100K<n<1M1 likes219 downloads11mo agoHugging Face04amphora /dasd-stage1-50k DASD stage1 - 50k length-filtered subset A 50,000-example subset of the stage1 (low-temperature) config of Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b. Columns are input / output; output is the verbatim gpt-oss-120b <think> reasoning trace. How it was built Started from stage1 (104,829 rows). Applied the Qwen3-4B-Instruct-2507 chat template and tokenized the full formatted conversation, then dropped every example over 65,536 tokens (the 64K training… See the full description on the dataset page: https://huggingface.co/datasets/amphora/dasd-stage1-50k.texttext-generation10K<n<100K0 likes190 downloads2mo agoHugging Face05fza1796262052 /DAS-0910-1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "observation.imu": { "dtype": "float32", "shape": [ 6 ], "names": [ "angular_velocity.x", "angular_velocity.y", "angular_velocity.z", "linear_acceleration.x", "linear_acceleration.y"… See the full description on the dataset page: https://huggingface.co/datasets/fza1796262052/DAS-0910-1.tabularrobotics1K<n<10K1 likes180 downloads10d agoHugging Face06stcoats /DASS2019_NLP Dataset Card for DASS2019_NLP This dataset contains audio and transcript content from DASS2019, the manually transcribed version of the Digital Archive of Southern Speech. It may be suitable for speech-related NLP processing, modelling, and fine-tuning tasks. Dataset Details DASS (Kretzschmar et al. 2012) comprises dialectological interviews with 64 informants conducted between 1968 and 1983; it is a subset of the larger Linguistic Atlas of the Gulf States (LAGS… See the full description on the dataset page: https://huggingface.co/datasets/stcoats/DASS2019_NLP.audioautomatic-speech-recognition10K<n<100K0 likes174 downloads6mo agoHugging Face07dasfas1132 /lelab-test_20260906_162149This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/dasfas1132/lelab-test_20260906_162149.tabularrobotics1K<n<10K0 likes172 downloads18d agoHugging Face08YanNeu /DASH-B DASH-B Object Hallucination Benchmark for Vision Language Models (VLMs) from the paper DASH: Detection and Assessment of Systematic Hallucinations of VLMs Model Evaluation | Citation Dataset The benchmark consists of 2682 images for a range of 70 different objects. The used query is "Can you see a object in this image. Please answer only with yes or no." 1341 of the images do not contain the corresponding object but trigger object hallucinations. They were retrieved… See the full description on the dataset page: https://huggingface.co/datasets/YanNeu/DASH-B.image1K<n<10K0 likes165 downloads1y agoHugging Face09optimum-benchmark /llm-perf-dashboardtextn<1K0 likes153 downloads2y agoHugging Face10Dasool /butterflies_and_moths_vqa Butterflies and Moths VQA Dataset Summary butterflies_and_moths_vqa is a visual question answering (VQA) dataset focused on butterflies and moths. It features tasks such as fine-grained species classification and ecological reasoning. The dataset is designed to benchmark Vision-Language Models (VLMs) for both image-based and text-only training approaches. Key Features Fine-Grained Classification (Type1): Questions requiring detailed species identification.… See the full description on the dataset page: https://huggingface.co/datasets/Dasool/butterflies_and_moths_vqa.imagevisual-question-answeringn<1K2 likes136 downloads2y agoHugging Face11daslab-testing /Apertus-8B-2509-microQAT-logitsThis dataset provides a small sample of TOP-K logits computed using swiss-ai/Apertus-8B-2509 on samples from Data Phase 5 of Apertus pre-training. Format This data represents documents packed into chuncks of 4096 tokens separated by EOS. The provided fields are as follows: input_ids: Input tokens. index: Positions of top-256 highest-probability next-token predictions for each token. exp_logits: Normalized probabilities of top-256 highest-probability next-token predictions for each… See the full description on the dataset page: https://huggingface.co/datasets/daslab-testing/Apertus-8B-2509-microQAT-logits.text-generation10K<n<100K0 likes120 downloads6mo agoHugging Face12fza1796262052 /DAS-0910This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "observation.imu": { "dtype": "float32", "shape": [ 6 ], "names": [ "angular_velocity.x", "angular_velocity.y", "angular_velocity.z", "linear_acceleration.x", "linear_acceleration.y"… See the full description on the dataset page: https://huggingface.co/datasets/fza1796262052/DAS-0910.tabularrobotics1K<n<10K1 likes120 downloads10d agoHugging Face13fza1796262052 /DAS-0910-2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "observation.imu": { "dtype": "float32", "shape": [ 6 ], "names": [ "angular_velocity.x", "angular_velocity.y", "angular_velocity.z", "linear_acceleration.x", "linear_acceleration.y"… See the full description on the dataset page: https://huggingface.co/datasets/fza1796262052/DAS-0910-2.tabularrobotics1K<n<10K0 likes113 downloads10d agoHugging Face14Dash00 /bc5cdr-disease-selectiontext1K<n<10K0 likes111 downloads8mo agoHugging Face15daspartho /urban_dictionarytext10K<n<100K9 likes107 downloads4y agoHugging Face16Dasool /KoMultiText KoMultiText: Korean Multi-task Dataset for Classifying Biased Speech Dataset Summary KoMultiText is a comprehensive Korean multi-task text dataset designed for classifying biased and harmful speech in online platforms. The dataset focuses on tasks such as Preference Detection, Profanity Identification, and Bias Classification across multiple domains, enabling state-of-the-art language models to perform multi-task learning for socially responsible AI applications.… See the full description on the dataset page: https://huggingface.co/datasets/Dasool/KoMultiText.tabulartext-classification10K<n<100K5 likes106 downloads8mo agoHugging Face17Yunis147 /trace-dashed-lineThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Yunis147/trace-dashed-line.tabularrobotics10K<n<100K0 likes95 downloads1mo agoHugging Face18daspartho /stable-diffusion-promptsSubset dataset of diffusiondb consisting of just unique prompts. Created this subset dataset for the Prompt Extend project. text1M<n<10M25 likes84 downloads3y agoHugging Face19dashuai2025 /mqtt_robot_demo_with_camThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "mqtt_robot", "total_episodes": 1, "total_frames": 1282, "total_tasks": 1, "total_videos": 0, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dashuai2025/mqtt_robot_demo_with_cam.tabularrobotics1K<n<10K0 likes71 downloads1y agoHugging Face20Dashboard-Integrity-Guard /test Dataset Card for "test" More Information needed image10K<n<100K1 likes66 downloads2y agoHugging Face21dassum /celebrity-identities Dataset Card for "celebrity-identities" More Information needed imagen<1K1 likes62 downloads3y agoHugging Face22ahmed-masry /DashboardQA DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards 🤗Dataset | 🖥️Code | 📄Paper The abstract of the paper states that: Dashboards are powerful visualization tools for data-driven decision-making, integrating multiple interactive views that allow users to explore, filter, and navigate data. Unlike static charts, dashboards support rich interactivity, which is essential for uncovering insights in real-world analytical workflows. However… See the full description on the dataset page: https://huggingface.co/datasets/ahmed-masry/DashboardQA.tabularn<1K0 likes56 downloads8mo agoHugging Face23Yunis147 /trace-dashed-line_20260820_141731This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Yunis147/trace-dashed-line_20260820_141731.tabularrobotics1K<n<10K0 likes56 downloads1mo agoHugging Face24Dashyash /indian_cuisine_datasetimage10K<n<100K0 likes55 downloads1y agoHugging Face25tsilva /gymrec__BreakoutNoFrameskip_dash_v4 BreakoutNoFrameskip-v4 Gameplay Dataset Gameplay recordings (collected by: breakout) from the Gymnasium environment BreakoutNoFrameskip-v4, captured using gymrec. Dataset Summary Stat Value Total frames 494,029 Episodes 50 Environment BreakoutNoFrameskip-v4 Backend Atari (ALE-py) Collector(s) breakout gymrec version(s) 0.1.0+23e91c8, 0.1.0+72aad18 Environment Configuration Setting Value Frameskip 1 Target FPS 30… See the full description on the dataset page: https://huggingface.co/datasets/tsilva/gymrec__BreakoutNoFrameskip_dash_v4.imagereinforcement-learning100K<n<1M0 likes50 downloads7mo agoHugging Face26DASHDarcy /GUI-Libra-81K-RLtextn<1K0 likes42 downloads26d agoHugging Face27Dasool /DC_inside_comments DC_inside_comments This dataset contains 110,000 raw comments collected from DC Inside. It is intended for unsupervised learning or pretraining purposes. Dataset Summary Data Type: Unlabeled raw comments Number of Examples: 110,000 Source: DC Inside Related Dataset For labeled data and multi-task annotated examples, please refer to the KoMultiText dataset. How to Load the Dataset from datasets import load_dataset # Load the unlabeled dataset… See the full description on the dataset page: https://huggingface.co/datasets/Dasool/DC_inside_comments.texttext-generation100K<n<1M0 likes41 downloads2y agoHugging Face28tsilva /gymrec__BreakoutNoFrameskip_dash_v4_stack4_unroll8_train_ready10K<n<100K0 likes41 downloads7mo agoHugging Face29Dash00 /ncbi-disease-selectiontext10K<n<100K0 likes40 downloads8mo agoHugging Face30davidkling /hf-coding-tools-dashboard-v2 HuggingFace AI Coding Tools Dashboard (Enhanced) Enhanced benchmark data from the HuggingFace AI Dashboard — includes query metadata (query_set, intent), run metadata (run_name, run_date), and freshness flags for stale references. This is the v2 enhanced dataset. The original dataset is at davidkling/hf-coding-tools-dashboard. Dataset Structure Split Description Rows results Enhanced results with query/run metadata and freshness flags 9146 queries… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-dashboard-v2.tabulartext-generation1K<n<10K1 likes39 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.