datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LSDIRdanish-dynaword
🧨 Danish Dynaword
Version
1.2.23 (Changelog)
Language
dan, dansk, Danish
License
Openly Licensed, See the respective dataset
Models
For model trained used this data see danish-foundation-models
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 7.40M
Number of tokens (Llama 3): 9.81B
Average document length in tokens (min, max): 1.33K (2, 19.46M)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-dynaword.yt-danish-public-v2FineWeb-Edu-10B-PMI-Filteredcurated-danbooru-2026
Curated Danbooru Streaming Dataset
A large-scale, high-performance curated dataset of ~330,000 (330K) high-quality anime illustrations designed for training Diffusion Transformers (DiT), Latent Diffusion Models (LDM), and text-to-image generative models focused on the anime domain.
This dataset is focused on specific curated characters and high-ranking artists using knowledge base lists (characters_list.txt and artists_list.txt). Prompt sequence lengths and bucket tiers (77… See the full description on the dataset page: https://huggingface.co/datasets/aipracticecafe/curated-danbooru-2026.danbooru-tags
List of Most Used Danbooru Tags
Contains a list of the most commonly used Danbooru tags, along with their usage statistics and metadata. I fetched them based on the following filters:
Order - Count
Is deprecated? - no
Hide Empty? - yes
Has Wiki - yes
Has artist - no
The dataset is available in the following formats:
tags.json
tags.jsonl
tags.parquet
danish-asr-unified
Danish ASR Unified Dataset
Unified Danish speech recognition dataset combining 7 sources (~3.5M samples, ~16k hours):
Source
Samples
Description
VoxPopuli
1,775,578
European Parliament recordings
ftspeech
995,677
Danish Parliament (Folketinget)
CoRal-v3 read_aloud
299,255
Read-aloud Danish speech
nst-da
182,605
NST Danish speech
CoRal-v3 conversation
147,249
Conversational Danish speech
nota
98,600
Danish broadcast media
Common Voice 17
3,484
Crowd-sourced… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified.urbansound8K(card and dataset copied from https://www.kaggle.com/datasets/chrisfilo/urbansound8k)
This dataset contains 8732 labeled sound excerpts (<=4s) of urban sounds from 10 classes: air_conditioner, car_horn, children_playing, dog_bark, drilling, enginge_idling, gun_shot, jackhammer, siren, and street_music. The classes are drawn from the urban sound taxonomy. For a detailed description of the dataset and how it was compiled please refer to our paper.All excerpts are taken from field recordings… See the full description on the dataset page: https://huggingface.co/datasets/danavery/urbansound8K.dancing-stick-figures
Dancing Stick Figures — v0.2
A small, fully-labelled synthetic video dataset for learning (and teaching) video diffusion on one consumer GPU.
1,340 clips · 6 s @ 20 fps · 128×128 RGBA · 482,400 frames · 134 text prompts × 10 seeds × 3 cameras ·
every frame carries the 3D skeleton, camera and G-buffer (depth, normals, part segmentation) that produced it.
Think of it as an MNIST for video generation: small enough that a 64² video diffusion model trains from scratch
in a few… See the full description on the dataset page: https://huggingface.co/datasets/sprited/dancing-stick-figures.norwegian-dynaword
🧨 Norwegian Dynaword
Version
0.0.18 (Changelog)
Language
Norwegian (no, nor), including Bokmål (nb, nob) and Nynorsk (nn, nno)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 4.47M
Number of tokens (Llama 3): 9.98B
Average document length in tokens (min… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dynaword.HERBench
HERBench: A Benchmark for Multi-Evidence Integration in Video Question Answering
A challenging benchmark for evaluating multi-evidence integration capabilities of vision-language models
🎉 HERBench has been accepted to CVPR 2026!
🆕 New: Lite-v2 config. We released a refreshed lite_v2 version of the
Lite split (1,971 questions / 68 videos) in which 9 of the 12 tasks were
regenerated and went through additional manual refinement for higher
quality, while TSO, SVA… See the full description on the dataset page: https://huggingface.co/datasets/DanBenAmi/HERBench.DanQing100M
100M Chinese image-text pairs | 12TB dataset | 2024-2025 web data
DanQing: An Up-to-Date Large-Scale Chinese Vision-Language Pre-training Dataset
Project Page | Paper | Code
Hengyu Shen∗, Tiancheng Gu∗, Bin Qin, Lan Wu, Yuling Wu, Shuo Tan, Zelong Sun, Jun Wang, Nan Wu, Xiang An, Weidong Cai, Ziyong Feng‡, Kaicheng Yang†
∗ Equal Contribution | ‡ Team Leader | † Project Leader
📣 News
[2026/01/16] ✨ We release the paper of DanQing.
[2026/01/15] 🔥 We release the… See the full description on the dataset page: https://huggingface.co/datasets/DeepGlint-AI/DanQing100M.danish-asr-unified-hviske-v5-tiny
danish-asr-unified — two-model labels and a quality manifest
Transcriptions, per-token confidences, and a per-row quality verdict for every
row of syvai/danish-asr-unified
(3,414,589 rows, 8 sources).
Two independently trained models labelled the whole corpus:
model
architecture
vocabulary
syvai/hviske-v5-tiny
encoder-decoder
16,384 BPE
3dio-ai/svale-110M
RNN-T (Parakeet)
44 characters
Both decode greedily. Shards mirror the source data/train-*.parquet by name… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified-hviske-v5-tiny.img-256-danbooru
Dataset Card for "img-256-danbooru"
More Information needed
AFRLA-assessor-instance-level-results
Assessors For Regression: Loss Analysis - Assessor Instance Level Results
Instance level results for assessors models trained on the AFRLA - Instance Level Results dataset.
At the moment of upload, results for XGBoost and linear regression models are available, with results from the former in 5 different seeds. Results are available for all 11 tasks described in the original dataset as well as for 6 different types of error (losses):
Loss name
Description… See the full description on the dataset page: https://huggingface.co/datasets/DaniFrame/AFRLA-assessor-instance-level-results.danish-asr-leaderboard
Open Danish ASR Leaderboard — Results
Benchmark results backing the Open Danish ASR Leaderboard — an open, reproducible comparison of Danish speech-to-text models.
Every model is transcribed and scored identically on the same five independent public Danish test sets, so the numbers compare directly: Word Error Rate (WER) and Character Error Rate (CER) — lower is better — plus speed. Open-weight Danish speech recognition models you can run yourself and hosted transcription APIs… See the full description on the dataset page: https://huggingface.co/datasets/RyeAI/danish-asr-leaderboard.LSDIR_rawdanbooru2025-metadata
🎨 Danbooru 2025 Metadata
Latest Post ID: 9,158,800
(as of Apr 16, 2025)
📁 About the DatasetThis dataset provides structured metadata for user-submitted images on Danbooru, a large-scale imageboard focused on anime-style artwork.
Scraping began on January 2, 2025, and the data are stored in Parquet format for efficient programmatic access.Compared to earlier versions, this snapshot includes:
More consistent tag history tracking
Better coverage of older or previously… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/danbooru2025-metadata.swedish-dynaword
🧨 Swedish Dynaword
Version
0.0.13 (Changelog)
Language
Swedish (sv, swe)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 547.06M
Number of tokens (Llama 3): 36.34B
Average document length in tokens (min, max): 66.42 (2, 8.14M)
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/swedish-dynaword.FineWeb-Edu-10B-Nouns-OnlyFineWeb-Edu-10B-Obfuscationcurated-danbooru-2026-512px-flux2-vaeaudioset_opus_24kbpschexpert
CheXpert
CheXpert is a large dataset of chest X-rays and competition for automated chest x-ray interpretation, which features uncertainty labels and radiologist-labeled reference standard evaluation sets.
https://stanfordmlgroup.github.io/competitions/chexpert/
Warning on AP/PA label
I could not find in the paper a mapping from the 0/1 label to AP/PA, so I assumed 0=AP and 1=PA. Looking at a few images this seems to be correct, but I'm not a radiologist.… See the full description on the dataset page: https://huggingface.co/datasets/danjacobellis/chexpert.danish-wit
Dataset Card for Danish WIT
Dataset Summary
Google presented the Wikipedia Image Text (WIT) dataset in July
2021, a dataset which contains
scraped images from Wikipedia along with their descriptions. WikiMedia released
WIT-Base in September
2021,
being a modified version of WIT where they have removed the images with empty
"reference descriptions", as well as removing images where a person's face covers more
than 10% of the image surface, along with inappropriate images… See the full description on the dataset page: https://huggingface.co/datasets/severo/danish-wit.ChartQA_small_preprocesseddutch-dynaword
🧨 Dutch Dynaword
Version
1.0.1 (Changelog)
Language
nld, Nederlands, Dutch
License
Openly Licensed, See the respective dataset
Models
For model trained used this data see danish-foundation-models
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 14.45M
Number of tokens (Llama 3): 37.89B
Average document length in tokens (min, max): 2.62K (2, 5.45M)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/dutch-dynaword.librusec
Dataset Card for "librusec"
More Information needed
icelandic-dynaword
🧨 Icelandic Dynaword
Version
0.0.15 (Changelog)
Language
Icelandic (is, isl)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 39.85M
Number of tokens (Llama 3): 2.67B
Average document length in tokens (min, max): 66.98 (3, 1.03M)
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dynaword.curated-danbooru-2026-256px-flux2-vae
