CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01danjacobellis /LSDIRimage10K<n<100K3 likes15k downloads2y agoHugging Face02danish-foundation-models /danish-dynaword 🧨 Danish Dynaword Version 1.2.23 (Changelog) Language dan, dansk, Danish License Openly Licensed, See the respective dataset Models For model trained used this data see danish-foundation-models Contact If you have question about this project please create an issue here Dataset Description Number of samples: 7.40M Number of tokens (Llama 3): 9.81B Average document length in tokens (min, max): 1.33K (2, 19.46M) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-dynaword.imagetext-generation10M<n<100M22 likes12k downloads25d agoHugging Face03overflowwwww /yt-danish-public-v2audioaudio-classification100K<n<1M0 likes2.7k downloads2y agoHugging Face04DanielGallagherIRE /FineWeb-Edu-10B-PMI-Filteredtext1M<n<10M0 likes2.6k downloads3mo agoHugging Face05aipracticecafe /curated-danbooru-2026 Curated Danbooru Streaming Dataset A large-scale, high-performance curated dataset of ~330,000 (330K) high-quality anime illustrations designed for training Diffusion Transformers (DiT), Latent Diffusion Models (LDM), and text-to-image generative models focused on the anime domain. This dataset is focused on specific curated characters and high-ranking artists using knowledge base lists (characters_list.txt and artists_list.txt). Prompt sequence lengths and bucket tiers (77… See the full description on the dataset page: https://huggingface.co/datasets/aipracticecafe/curated-danbooru-2026.tabulartext-to-image100K<n<1M9 likes2.3k downloads10d agoHugging Face06qdlabs /danbooru-tags List of Most Used Danbooru Tags Contains a list of the most commonly used Danbooru tags, along with their usage statistics and metadata. I fetched them based on the following filters: Order - Count Is deprecated? - no Hide Empty? - yes Has Wiki - yes Has artist - no The dataset is available in the following formats: tags.json tags.jsonl tags.parquet tabular100K<n<1M15 likes2.2k downloads1y agoHugging Face07syvai /danish-asr-unified Danish ASR Unified Dataset Unified Danish speech recognition dataset combining 7 sources (~3.5M samples, ~16k hours): Source Samples Description VoxPopuli 1,775,578 European Parliament recordings ftspeech 995,677 Danish Parliament (Folketinget) CoRal-v3 read_aloud 299,255 Read-aloud Danish speech nst-da 182,605 NST Danish speech CoRal-v3 conversation 147,249 Conversational Danish speech nota 98,600 Danish broadcast media Common Voice 17 3,484 Crowd-sourced… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified.audioautomatic-speech-recognition1M<n<10M5 likes2.2k downloads2mo agoHugging Face08danavery /urbansound8K(card and dataset copied from https://www.kaggle.com/datasets/chrisfilo/urbansound8k) This dataset contains 8732 labeled sound excerpts (<=4s) of urban sounds from 10 classes: air_conditioner, car_horn, children_playing, dog_bark, drilling, enginge_idling, gun_shot, jackhammer, siren, and street_music. The classes are drawn from the urban sound taxonomy. For a detailed description of the dataset and how it was compiled please refer to our paper.All excerpts are taken from field recordings… See the full description on the dataset page: https://huggingface.co/datasets/danavery/urbansound8K.audioaudio-classification1K<n<10K10 likes2.1k downloads3y agoHugging Face09sprited /dancing-stick-figures Dancing Stick Figures — v0.2 A small, fully-labelled synthetic video dataset for learning (and teaching) video diffusion on one consumer GPU. 1,340 clips · 6 s @ 20 fps · 128×128 RGBA · 482,400 frames · 134 text prompts × 10 seeds × 3 cameras · every frame carries the 3D skeleton, camera and G-buffer (depth, normals, part segmentation) that produced it. Think of it as an MNIST for video generation: small enough that a 64² video diffusion model trains from scratch in a few… See the full description on the dataset page: https://huggingface.co/datasets/sprited/dancing-stick-figures.imagetext-to-video100K<n<1M2 likes2.1k downloads1mo agoHugging Face10danish-foundation-models /norwegian-dynaword 🧨 Norwegian Dynaword Version 0.0.18 (Changelog) Language Norwegian (no, nor), including Bokmål (nb, nob) and Nynorsk (nn, nno) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 4.47M Number of tokens (Llama 3): 9.98B Average document length in tokens (min… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dynaword.imagetext-generation10M<n<100M7 likes2.1k downloads19d agoHugging Face11DanBenAmi /HERBench HERBench: A Benchmark for Multi-Evidence Integration in Video Question Answering A challenging benchmark for evaluating multi-evidence integration capabilities of vision-language models 🎉 HERBench has been accepted to CVPR 2026! 🆕 New: Lite-v2 config. We released a refreshed lite_v2 version of the Lite split (1,971 questions / 68 videos) in which 9 of the 12 tasks were regenerated and went through additional manual refinement for higher quality, while TSO, SVA… See the full description on the dataset page: https://huggingface.co/datasets/DanBenAmi/HERBench.tabularvisual-question-answering10K<n<100K3 likes2k downloads4mo agoHugging Face12DeepGlint-AI /DanQing100M 100M Chinese image-text pairs | 12TB dataset | 2024-2025 web data DanQing: An Up-to-Date Large-Scale Chinese Vision-Language Pre-training Dataset Project Page | Paper | Code Hengyu Shen∗, Tiancheng Gu∗, Bin Qin, Lan Wu, Yuling Wu, Shuo Tan, Zelong Sun, Jun Wang, Nan Wu, Xiang An, Weidong Cai, Ziyong Feng‡, Kaicheng Yang† ∗ Equal Contribution | ‡ Team Leader | † Project Leader 📣 News [2026/01/16] ✨ We release the paper of DanQing. [2026/01/15] 🔥 We release the… See the full description on the dataset page: https://huggingface.co/datasets/DeepGlint-AI/DanQing100M.imagezero-shot-image-classification10M<n<100M52 likes1.9k downloads6mo agoHugging Face13syvai /danish-asr-unified-hviske-v5-tinygated danish-asr-unified — two-model labels and a quality manifest Transcriptions, per-token confidences, and a per-row quality verdict for every row of syvai/danish-asr-unified (3,414,589 rows, 8 sources). Two independently trained models labelled the whole corpus: model architecture vocabulary syvai/hviske-v5-tiny encoder-decoder 16,384 BPE 3dio-ai/svale-110M RNN-T (Parakeet) 44 characters Both decode greedily. Shards mirror the source data/train-*.parquet by name… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified-hviske-v5-tiny.tabularautomatic-speech-recognition1M<n<10M0 likes1.9k downloads10d agoHugging Face14woctordho /img-256-danbooru Dataset Card for "img-256-danbooru" More Information needed image100K<n<1M0 likes1.9k downloads4y agoHugging Face15DaniFrame /AFRLA-assessor-instance-level-results Assessors For Regression: Loss Analysis - Assessor Instance Level Results Instance level results for assessors models trained on the AFRLA - Instance Level Results dataset. At the moment of upload, results for XGBoost and linear regression models are available, with results from the former in 5 different seeds. Results are available for all 11 tasks described in the original dataset as well as for 6 different types of error (losses): Loss name Description… See the full description on the dataset page: https://huggingface.co/datasets/DaniFrame/AFRLA-assessor-instance-level-results.tabular10M<n<100M0 likes1.8k downloads2y agoHugging Face16RyeAI /danish-asr-leaderboard Open Danish ASR Leaderboard — Results Benchmark results backing the Open Danish ASR Leaderboard — an open, reproducible comparison of Danish speech-to-text models. Every model is transcribed and scored identically on the same five independent public Danish test sets, so the numbers compare directly: Word Error Rate (WER) and Character Error Rate (CER) — lower is better — plus speed. Open-weight Danish speech recognition models you can run yourself and hosted transcription APIs… See the full description on the dataset page: https://huggingface.co/datasets/RyeAI/danish-asr-leaderboard.tabularautomatic-speech-recognition1M<n<10M4 likes1.8k downloads14h agoHugging Face17danjacobellis /LSDIR_rawimage10K<n<100K1 likes1.6k downloads2y agoHugging Face18trojblue /danbooru2025-metadata 🎨 Danbooru 2025 Metadata Latest Post ID: 9,158,800 (as of Apr 16, 2025) 📁 About the DatasetThis dataset provides structured metadata for user-submitted images on Danbooru, a large-scale imageboard focused on anime-style artwork. Scraping began on January 2, 2025, and the data are stored in Parquet format for efficient programmatic access.Compared to earlier versions, this snapshot includes: More consistent tag history tracking Better coverage of older or previously… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/danbooru2025-metadata.imagetext-to-image1M<n<10M38 likes1.6k downloads1y agoHugging Face19danish-foundation-models /swedish-dynaword 🧨 Swedish Dynaword Version 0.0.13 (Changelog) Language Swedish (sv, swe) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 547.06M Number of tokens (Llama 3): 36.34B Average document length in tokens (min, max): 66.42 (2, 8.14M) Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/swedish-dynaword.imagetext-generation1B<n<10B3 likes1.6k downloads17d agoHugging Face20DanielGallagherIRE /FineWeb-Edu-10B-Nouns-Onlytext1M<n<10M0 likes1.6k downloads3mo agoHugging Face21DanielGallagherIRE /FineWeb-Edu-10B-Obfuscationtext10M<n<100M0 likes1.3k downloads4mo agoHugging Face22aipracticecafe /curated-danbooru-2026-512px-flux2-vaetabular100K<n<1M0 likes1.3k downloads26d agoHugging Face23danjacobellis /audioset_opus_24kbpsaudio1M<n<10M1 likes1.2k downloads2y agoHugging Face24danjacobellis /chexpert CheXpert CheXpert is a large dataset of chest X-rays and competition for automated chest x-ray interpretation, which features uncertainty labels and radiologist-labeled reference standard evaluation sets. https://stanfordmlgroup.github.io/competitions/chexpert/ Warning on AP/PA label I could not find in the paper a mapping from the 0/1 label to AP/PA, so I assumed 0=AP and 1=PA. Looking at a few images this seems to be correct, but I'm not a radiologist.… See the full description on the dataset page: https://huggingface.co/datasets/danjacobellis/chexpert.imageimage-classification100K<n<1M23 likes1.1k downloads2y agoHugging Face25severo /danish-wit Dataset Card for Danish WIT Dataset Summary Google presented the Wikipedia Image Text (WIT) dataset in July 2021, a dataset which contains scraped images from Wikipedia along with their descriptions. WikiMedia released WIT-Base in September 2021, being a modified version of WIT where they have removed the images with empty "reference descriptions", as well as removing images where a person's face covers more than 10% of the image surface, along with inappropriate images… See the full description on the dataset page: https://huggingface.co/datasets/severo/danish-wit.imageimage-to-text100K<n<1M0 likes1k downloads4y agoHugging Face26DanhVuiVe /ChartQA_small_preprocessedimage1K<n<10K0 likes945 downloads2y agoHugging Face27danish-foundation-models /dutch-dynaword 🧨 Dutch Dynaword Version 1.0.1 (Changelog) Language nld, Nederlands, Dutch License Openly Licensed, See the respective dataset Models For model trained used this data see danish-foundation-models Contact If you have question about this project please create an issue here Dataset Description Number of samples: 14.45M Number of tokens (Llama 3): 37.89B Average document length in tokens (min, max): 2.62K (2, 5.45M) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/dutch-dynaword.imagetext-generation10M<n<100M3 likes926 downloads16d agoHugging Face28danasone /librusec Dataset Card for "librusec" More Information needed text1K<n<10K1 likes863 downloads3y agoHugging Face29danish-foundation-models /icelandic-dynaword 🧨 Icelandic Dynaword Version 0.0.15 (Changelog) Language Icelandic (is, isl) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 39.85M Number of tokens (Llama 3): 2.67B Average document length in tokens (min, max): 66.98 (3, 1.03M) Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dynaword.imagetext-generation100M<n<1B4 likes860 downloads19d agoHugging Face30aipracticecafe-mirror /curated-danbooru-2026-256px-flux2-vaetabular100K<n<1M0 likes820 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.