datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
danish-dynaword
🧨 Danish Dynaword
Version
1.2.23 (Changelog)
Language
dan, dansk, Danish
License
Openly Licensed, See the respective dataset
Models
For model trained used this data see danish-foundation-models
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 7.40M
Number of tokens (Llama 3): 9.81B
Average document length in tokens (min, max): 1.33K (2, 19.46M)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-dynaword.DanbooruwildcardsThis is a set of wildcards for danbooru tags.
Artist:Prompts for random artist styles, covering approximately 0.6M different artists.Please select the appropriate version of the collection, ranging from 128 to 5000, based on the model's capabilities.The full version is not recommended for use as it includes too many artists with only one image on danbooru or other websites. Almost no model can generate a style that corresponds to these artists .
Characters:"Characters" is a set of wildcards… See the full description on the dataset page: https://huggingface.co/datasets/X779/Danbooruwildcards.norwegian-dynaword
🧨 Norwegian Dynaword
Version
0.0.18 (Changelog)
Language
Norwegian (no, nor), including Bokmål (nb, nob) and Nynorsk (nn, nno)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 4.47M
Number of tokens (Llama 3): 9.98B
Average document length in tokens (min… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dynaword.swedish-dynaword
🧨 Swedish Dynaword
Version
0.0.13 (Changelog)
Language
Swedish (sv, swe)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 547.06M
Number of tokens (Llama 3): 36.34B
Average document length in tokens (min, max): 66.42 (2, 8.14M)
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/swedish-dynaword.dutch-dynaword
🧨 Dutch Dynaword
Version
1.0.1 (Changelog)
Language
nld, Nederlands, Dutch
License
Openly Licensed, See the respective dataset
Models
For model trained used this data see danish-foundation-models
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 14.45M
Number of tokens (Llama 3): 37.89B
Average document length in tokens (min, max): 2.62K (2, 5.45M)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/dutch-dynaword.icelandic-dynaword
🧨 Icelandic Dynaword
Version
0.0.15 (Changelog)
Language
Icelandic (is, isl)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 39.85M
Number of tokens (Llama 3): 2.67B
Average document length in tokens (min, max): 66.98 (3, 1.03M)
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dynaword.danbooru-wiki-2024
danbooru-wiki-2024
About
Wiki pages about the danbooru tags on danbooru.donmai.us. The wiki contains the description of each tag and matching to pixiv tags.
Usage
from datasets import load_dataset
ds = load_dataset(
"isek-ai/danbooru-wiki-2024",
# revision="202408-at20240906", # optional
split="train",
)
The revision name is as same as isek-ai/danbooru-tags-2024's.
[!WARNING]
Note:
This dataset would be irreguraly updated, if you want to use the same… See the full description on the dataset page: https://huggingface.co/datasets/isek-ai/danbooru-wiki-2024.faroese-dynaword
🧨 Faroese Dynaword
Version
0.0.7 (Changelog)
Language
Faroese (fo, fao)
License
Openly Licensed, See the respective dataset
Models
Currently there are no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 405.81K
Number of tokens (Llama 3): 45.40M
Average document length in tokens (min, max): 111.87 (2, 109.50K)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/faroese-dynaword.public-domain-poetry
Overview
This dataset is a collection of approximately 38,500 poems from https://www.public-domain-poetry.com/.
Language
The language of this dataset is English.
License
All data in this dataset is public domain, which means you should be able to use it for anything you want, as long as you aren't breaking any law in the process of doing so.
danbooru-tags-2024
danbooru-tags-2024
from datasets import load_dataset
ds = load_dataset(
"isek-ai/danbooru-tags-2024",
# revision="202412-at20250122", # optional
split="train",
)
Last updated: since 2005 to 2024/12/31, collected at 2025/01/22
danish-gigaword
Danish Gigaword Corpus
Version: 1.0.0
License: See the respective dataset
Dataset Summary
The Danish Gigaword Corpus contains text spanning several domains and forms. This version does not include the sections containing tweets ("General Discussions" and "Parliament Elections"), "danavis", "Common Crawl" and "OpenSubtitles" due to potential privacy, quality and copyright concerns.
Loading the dataset
from datasets import load_dataset
name =… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-gigaword.GeneratingQuestions
HVU_QA
HVU_QA is an open-source Vietnamese Question-Context-Answer (QCA) corpus, accompanied by supporting tools, created to facilitate the development of FAQ-style question generation and question answering systems, particularly for low-resource language settings. The dataset was developed by a research team at Hung Vuong University, Phu Tho, Vietnam, led by Dr. Ha Nguyen, Deputy Head of the Department of Engineering Technology. HVU_QA was constructed using a fully automated… See the full description on the dataset page: https://huggingface.co/datasets/DANGDOCAO/GeneratingQuestions.Open-Router-API-Pricing-Analysis
OpenRouter API Pricing Analysis Dataset
Overview
This dataset provides a point-in-time capture of pricing and parameters for LLMs available through the OpenRouter API for inference.
Contents
Raw Data (raw/)
Contains the original data extracted from the OpenRouter API, including:
Model pricing (input/output token costs)
Model parameters and specifications
Computed fields such as output/input token price ratios
Enhanced Data (hf-enhanced/)… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Open-Router-API-Pricing-Analysis.HundredCV-Chat
百人对话数据集
HundredCV-Chat: A Dataset of Daily Chatting Developed on HundredCVs
简介
本项目提出一个全新的中文多轮对话数据集(HundredCV-Chat),该数据集由 100 位青年的简历数据集 HundredCVs 开发而来,共包含 24,750 组日常闲聊对话数据。
数据集具有如下特点:
自动化标注:HundredCV-Chat 中的对话均由 Deepseek-V3 大模型生成,不涉及任何人工标注,因此同时保证了大规模数据量和低成本优势。
多样性话题:HundredCV-Chat 中的对话话题涵盖了校园生活、工作经验、兴趣爱好、生活琐事等多个方面,与真实生活联系紧密,尤其适用于开发年轻化应用。
高质量对话:利用 Deepseek 强大的生成能力和全面的知识,HundredCV-Chat 的对话内容在流畅度、拟人性、多样性方面均显著优于现有的开源对话数据集。
数据样例
HundredCV-Chat 含有 24… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/HundredCV-Chat.danbooru-2408-blind-captions
Danbooru 2408 Blind Captions
from datasets import load_dataset
ds = load_dataset(
"dartags/danbooru-2408-blind-captions",
split="train",
)
openresearcher-sft-deep-research-cleaned
OpenResearcher SFT DeepResearch — Parquet Mirror
This is a re-hosted copy of the tool-reasoning SFT deep-research dataset by Aman Priyanshu, itself a cleaned/restructured version of the OpenResearcher Dataset from TIGER-AI-Lab.
Why this repo exists: the source wasn't laid out as ready-to-download Parquet files. This mirror simply stores the data as plain seed_*.parquet files so you can grab the whole dataset or a single segment easily. No changes were made to the content — all… See the full description on the dataset page: https://huggingface.co/datasets/DanielTobi0/openresearcher-sft-deep-research-cleaned.danbooru-wiki-2026
danbooru-wiki-2026-04-28
About
Wiki pages about the danbooru tags on danbooru.donmai.us. The wiki contains the description of each tag.
This dataset was collected from the wiki data using danbooru API and then merged with tags data (category_name & post_count). No entries were removed, no matter how many posts they have. The resulting set was manually filtered. Originally some of the lists, tag_group:, about:, help:, howto:, etc, pages were labled as "general"… See the full description on the dataset page: https://huggingface.co/datasets/kierarkia/danbooru-wiki-2026.danish-tool-dialogues-v9
danish-tool-dialogues-v1
Danish multi-turn tool-use conversations with reasoning, translated from the
Glaive subset of
Nanbeige/ToolMind
(Apache-2.0) by scripts/translate_toolmind_da.py.
Complements danish-tool-calls-v1, which is single-turn and synthetic. Here
the conversations run several turns, tool results are fed back, and the
assistant reasons before calling.
split
rows
train
34,168
eval_seen_tools
698
eval_unseen_tools
768
eval_seen_sym
752… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-tool-dialogues-v9.pandora-tool-calling
Pandora Tool Calling
A tool-calling dataset for Supervised fine-tuning of the Pandora Large Language Model (LLM).
The dataset is based on the glaiveai/glaive-function-calling-v2 dataset.
Copyright and license
Copyright (c) 2024, Danilo Peixoto Ferreira. All rights reserved.
Project developed under a BSD-3-Clause license.
norwegian-dyna-instruct
🧨 Norwegian dyna-instruct
Version
0.1.0 (changelog)
Languages
Norwegian Bokmål (nob), Norwegian Nynorsk (nno), and English (eng) translation input
License
Mixed open licenses; see the table below
Sources
Five datasets (source cards)
Dataset Description
Number of samples: 14.40K
Number of tokens (Llama 3): 6.27M
Average conversation length in tokens (min, max): 435.63 (4, 8.92K)
Average number of turns (min, max): 2.13 (2, 3)… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dyna-instruct.faroese-dyna-instruct
🧨 Faroese dyna-instruct
Version
0.1.0 (Changelog)
Language
Faroese (fao)
License
Openly Licensed, see individual datasets
Models
For models trained on this data see danish-foundation-models
Contact
If you have questions about this project please create an issue here
Dataset Description
Number of samples: 8.61K
Number of tokens (Llama 3): 2.64M
Average conversation length in tokens (min, max): 306.67 (98, 1.24K)
Average number of… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/faroese-dyna-instruct.danbooru-tags-2023
danbooru-tags-2023
A dataset of danbooru tags.
Dataset information
Generated using danbooru and safebooru API.
The dataset was created with the following conditions:
Subset name
all
safe
API Endpoint
https://danbooru.donmai.us
https://safebooru.donmai.us
Date
2005-01-01..2023-12-31
2005-01-01..2023-12-31
Score
>0
>0
Rating
g,s,q,e
g
Filetype
png,jpg,webppng,jpg,webp
Size (number of rows)
6,574,149
1,387,371
Usage
pip install datasets… See the full description on the dataset page: https://huggingface.co/datasets/isek-ai/danbooru-tags-2023.RetailBanking-Conversations
Dataset Description
RetailBanking-Conversations is a synthetic dataset designed to train and evaluate language models in the retail banking domain, it has been created using the open source library wizardSdata that eable the creation of synthetic datasets in any field.
The dataset contains 320 realistic conversations, across 160 unique financial profiles and 10 key retail banking topics, between financial advisors and clients, covering 10 main categories of banking products and… See the full description on the dataset page: https://huggingface.co/datasets/danystar/RetailBanking-Conversations.danish-extraction-v1
danish-extraction-v1
Danish information-extraction rows over real prose, where the schema is
proposed per passage rather than fixed. Built from
danish-foundation-models/danish-dynaword
by scripts/gen_extraction_da.py.
Each source passage got its own field set: an LLM proposed 3-6 fields for that
text without seeing any values, then filled them in a separate turn. Roughly a
quarter of proposed fields come back empty, which are genuine abstention
targets rather than annotation… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-extraction-v1.ifeval-da
IFEval-da
This dataset is a translation of the English IFEval dataset,
which was published in this paper and contains 541 prompts,
each with a combination of one or more of 25 different constraints. The dataset was professionally
translated and localised by expert native speakers.
Dataset Details
Translated by: Rasmus Larsen (rasmus.larsen@alexandra.dk), Nathalie Hau Sørensen (naha@hum.ku.dk) and Kenneth Enevoldsen (kenneth.enevoldsen@cas.au.dk)
Funded by: Danish… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/ifeval-da.danish-tool-dialogues-v6
danish-tool-dialogues-v1
Danish multi-turn tool-use conversations with reasoning, translated from the
Glaive subset of
Nanbeige/ToolMind
(Apache-2.0) by scripts/translate_toolmind_da.py.
Complements danish-tool-calls-v1, which is single-turn and synthetic. Here
the conversations run several turns, tool results are fed back, and the
assistant reasons before calling.
split
rows
train
17,667
eval_seen_tools
722
eval_unseen_tools
779
933 distinct tools; 59 names… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-tool-dialogues-v6.danbooru-tags-2024
danbooru-tags-2024
from datasets import load_dataset
ds = load_dataset(
"isek-ai/danbooru-tags-2024",
# revision="202412-at20250122", # optional
split="train",
)
Last updated: since 2005 to 2024/12/31, collected at 2025/01/22
danish-tool-dialogues-v7
danish-tool-dialogues-v1
Danish multi-turn tool-use conversations with reasoning, translated from the
Glaive subset of
Nanbeige/ToolMind
(Apache-2.0) by scripts/translate_toolmind_da.py.
Complements danish-tool-calls-v1, which is single-turn and synthetic. Here
the conversations run several turns, tool results are fed back, and the
assistant reasons before calling.
split
rows
train
17,160
eval_seen_tools
701
eval_unseen_tools
768
925 distinct tools; 59 names… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-tool-dialogues-v7.danish-tool-dialogues-v4
danish-tool-dialogues-v1
Danish multi-turn tool-use conversations with reasoning, translated from the
Glaive subset of
Nanbeige/ToolMind
(Apache-2.0) by scripts/translate_toolmind_da.py.
Complements danish-tool-calls-v1, which is single-turn and synthetic. Here
the conversations run several turns, tool results are fed back, and the
assistant reasons before calling.
split
rows
train
17,598
eval_seen_tools
762
eval_unseen_tools
772
932 distinct tools; 59 names… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-tool-dialogues-v4.danbooru-tags-2016-2023
danbooru-tags-2016-2023
A dataset of danbooru tags.
Dataset information
Generated using danbooru and safebooru API.
The dataset was created with the following conditions:
Subset name
all
safe
API Endpoint
https://danbooru.donmai.us
https://safebooru.donmai.us
Date
2016-01-01..2023-12-31
2016-01-01..2023-12-31
Score
>0
>0
Rating
g,s,q,e
g
Filetype
png,jpg,webppng,jpg,webp
Size (number of rows)
4,601,557
1,186,490
Usage
pip install… See the full description on the dataset page: https://huggingface.co/datasets/isek-ai/danbooru-tags-2016-2023.
