CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01danish-foundation-models /danish-dynaword 🧨 Danish Dynaword Version 1.2.23 (Changelog) Language dan, dansk, Danish License Openly Licensed, See the respective dataset Models For model trained used this data see danish-foundation-models Contact If you have question about this project please create an issue here Dataset Description Number of samples: 7.40M Number of tokens (Llama 3): 9.81B Average document length in tokens (min, max): 1.33K (2, 19.46M) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-dynaword.imagetext-generation10M<n<100M22 likes11k downloads21d agoHugging Face02danish-foundation-models /norwegian-dynaword 🧨 Norwegian Dynaword Version 0.0.18 (Changelog) Language Norwegian (no, nor), including Bokmål (nb, nob) and Nynorsk (nn, nno) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 4.47M Number of tokens (Llama 3): 9.98B Average document length in tokens (min… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dynaword.imagetext-generation10M<n<100M7 likes1.9k downloads15d agoHugging Face03danish-foundation-models /swedish-dynaword 🧨 Swedish Dynaword Version 0.0.13 (Changelog) Language Swedish (sv, swe) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 547.06M Number of tokens (Llama 3): 36.34B Average document length in tokens (min, max): 66.42 (2, 8.14M) Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/swedish-dynaword.imagetext-generation1B<n<10B3 likes1.4k downloads13d agoHugging Face04danish-foundation-models /dutch-dynaword 🧨 Dutch Dynaword Version 1.0.1 (Changelog) Language nld, Nederlands, Dutch License Openly Licensed, See the respective dataset Models For model trained used this data see danish-foundation-models Contact If you have question about this project please create an issue here Dataset Description Number of samples: 14.45M Number of tokens (Llama 3): 37.89B Average document length in tokens (min, max): 2.62K (2, 5.45M) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/dutch-dynaword.imagetext-generation10M<n<100M3 likes824 downloads12d agoHugging Face05danish-foundation-models /icelandic-dynaword 🧨 Icelandic Dynaword Version 0.0.15 (Changelog) Language Icelandic (is, isl) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 39.85M Number of tokens (Llama 3): 2.67B Average document length in tokens (min, max): 66.98 (3, 1.03M) Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dynaword.imagetext-generation100M<n<1B4 likes806 downloads15d agoHugging Face06danish-foundation-models /faroese-dynaword 🧨 Faroese Dynaword Version 0.0.7 (Changelog) Language Faroese (fo, fao) License Openly Licensed, See the respective dataset Models Currently there are no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 405.81K Number of tokens (Llama 3): 45.40M Average document length in tokens (min, max): 111.87 (2, 109.50K) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/faroese-dynaword.imagetext-generation1M<n<10M3 likes516 downloads6d agoHugging Face07danish-foundation-models /danish-gigaword Danish Gigaword Corpus Version: 1.0.0 License: See the respective dataset Dataset Summary The Danish Gigaword Corpus contains text spanning several domains and forms. This version does not include the sections containing tweets ("General Discussions" and "Parliament Elections"), "danavis", "Common Crawl" and "OpenSubtitles" due to potential privacy, quality and copyright concerns. Loading the dataset from datasets import load_dataset name =… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-gigaword.texttext-generation100K<n<1M9 likes468 downloads2y agoHugging Face08danish-foundation-models /ifeval-da IFEval-da This dataset is a translation of the English IFEval dataset, which was published in this paper and contains 541 prompts, each with a combination of one or more of 25 different constraints. The dataset was professionally translated and localised by expert native speakers. Dataset Details Translated by: Rasmus Larsen (rasmus.larsen@alexandra.dk), Nathalie Hau Sørensen (naha@hum.ku.dk) and Kenneth Enevoldsen (kenneth.enevoldsen@cas.au.dk) Funded by: Danish… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/ifeval-da.texttext-generationn<1K1 likes231 downloads7mo agoHugging Face09danish-foundation-models /norwegian-dyna-instruct 🧨 Norwegian dyna-instruct Version 0.1.0 (changelog) Languages Norwegian Bokmål (nob), Norwegian Nynorsk (nno), and English (eng) translation input License Mixed open licenses; see the table below Sources Five datasets (source cards) Dataset Description Number of samples: 14.40K Number of tokens (Llama 3): 6.27M Average conversation length in tokens (min, max): 435.63 (4, 8.92K) Average number of turns (min, max): 2.13 (2, 3)… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dyna-instruct.imagequestion-answering10K<n<100K0 likes230 downloads15d agoHugging Face10danish-foundation-models /faroese-dyna-instruct 🧨 Faroese dyna-instruct Version 0.1.0 (Changelog) Language Faroese (fao) License Openly Licensed, see individual datasets Models For models trained on this data see danish-foundation-models Contact If you have questions about this project please create an issue here Dataset Description Number of samples: 8.61K Number of tokens (Llama 3): 2.64M Average conversation length in tokens (min, max): 306.67 (98, 1.24K) Average number of… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/faroese-dyna-instruct.texttext-generation10K<n<100K1 likes217 downloads20d agoHugging Face11danish-foundation-models /icelandic-dyna-instruct 🧨 Icelandic dyna-instruct Version 0.1.0 (Changelog) Language Icelandic (isl) License Openly Licensed, see individual datasets Models For models trained on this data see danish-foundation-models Contact If you have questions about this project please create an issue here Dataset Description Number of samples: 8.11K Number of tokens (Llama 3): 7.09M Average conversation length in tokens (min, max): 874.89 (182, 1.39K) Average number… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dyna-instruct.texttext-generation10K<n<100K1 likes158 downloads20d agoHugging Face12danish-foundation-models /nasjonalt-vitenarkiv Nasjonalt vitenarkiv Open-access documents from NVA (Nasjonalt vitenarkiv), the joint national repository where Norwegian research institutions publish their output: master's and PhD theses, journal articles, and technical and research reports. Subjects span the disciplines - marine science, forestry, archaeology, education, public health, engineering - and most documents are recent. Each row is one PDF: the original file exactly as published, the text extracted from it, and the… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/nasjonalt-vitenarkiv.documenttext-generationn<1K1 likes83 downloads2mo agoHugging Face13bkai-foundation-models /vietnamese-roleplay-realm 🇻🇳 Vietnamese Role-play Realm Dataset This is a dataset of GPT-generated Vietnamese characters made to increase the ability of open-source language models to role-play. It contains 446 characters generated by GPT-3.5 Each character will have 20 topics generated by ChatGPT. And each topic will have a conversation corresponding with it In 446 characters, there are 400 general characters and 46 Vietnamese characters. To construct this dataset, we follow a four-step process:… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/vietnamese-roleplay-realm.text-generation3 likes54 downloads3y agoHugging Face14danish-foundation-models /laerebogengated Lærebogen An instruction-following dataset for Danish. This dataset features 5 million examples of multi-turn conversations in Danish, designed to train instruction-following models, with a commercially usable license. Dataset Structure All examples in the dataset are structured as follows: { "messages": [ { "role": "user", "content": "(...)" }, { "role": "assistant", "content": "(...)" }, { "role": "user", "content": "(...)" }, (...) { "role":… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/laerebogen.texttext-generation1M<n<10M1 likes28 downloads6mo agoHugging Face15danish-foundation-models /dfm-dyna-instructgated 🧨 DFM dyna-instruct Version 0.1.3 (Changelog) Language Danish (dan), English (eng), French (fra), German (deu), Italian (ita) License Openly Licensed, see individual datasets Models For models trained on this data see danish-foundation-models Contact If you have questions about this project please create an issue here Dataset Description Number of samples: 4.40M Number of tokens (Llama 3): 2.85B Average conversation length in tokens… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/dfm-dyna-instruct.imagetext-generation1M<n<10M4 likes24 downloads4mo agoHugging Face16fineset-io /time-series-foundation-models-papers Time Series Foundation Models Papers — FineSet A research-paper dataset on Time Series Foundation Models Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-19. It is not auto-updated. Research on Time Series Foundation Models Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/time-series-foundation-models-papers.tabulartext-classificationn<1K0 likes19 downloads3mo agoHugging Face17danish-foundation-models /ai-arenaen-conversationsgated AI Arenaen Conversations A large dataset of conversations from AI-Arenaen, the Danish subset of the compar:IA platform. Origin of the data: what is AI-Arenaen? The conversations are collected using AI-Arenaen, the Danish entry point to the compar:IA platform, which is a Conversational AI comparison tool (a "chatbot arena"), developed within the French Ministry of Culture and adapted for Danish users by Danish Foundation Models and The ministry of digital affair.… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/ai-arenaen-conversations.tabulartext-generation1K<n<10K1 likes13 downloads4mo agoHugging Face18foundation-models /eval-runs eval-runs Evaluation run artifacts from τ2-bench simulations. Layout tau2-bench/ gpt-4o-mini/ with-patch/ # Retail policy includes agent-lens write guardrail without-patch/ # Baseline τ2-bench retail policy (no guardrail) other-models/ with-patch/ without-patch/ # Smoke / scaffold runs on non–GPT-4o-mini models Patch = the retail policy was augmented with the agent-lens retail write guardrail block (write-action discipline).… See the full description on the dataset page: https://huggingface.co/datasets/foundation-models/eval-runs.text-generation0 likes12 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.