CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01danish-foundation-models /danish-dynaword 🧨 Danish Dynaword Version 1.2.23 (Changelog) Language dan, dansk, Danish License Openly Licensed, See the respective dataset Models For model trained used this data see danish-foundation-models Contact If you have question about this project please create an issue here Dataset Description Number of samples: 7.40M Number of tokens (Llama 3): 9.81B Average document length in tokens (min, max): 1.33K (2, 19.46M) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-dynaword.imagetext-generation10M<n<100M22 likes11k downloads22d agoHugging Face02danish-foundation-models /norwegian-dynaword 🧨 Norwegian Dynaword Version 0.0.18 (Changelog) Language Norwegian (no, nor), including Bokmål (nb, nob) and Nynorsk (nn, nno) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 4.47M Number of tokens (Llama 3): 9.98B Average document length in tokens (min… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dynaword.imagetext-generation10M<n<100M7 likes2k downloads16d agoHugging Face03danish-foundation-models /swedish-dynaword 🧨 Swedish Dynaword Version 0.0.13 (Changelog) Language Swedish (sv, swe) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 547.06M Number of tokens (Llama 3): 36.34B Average document length in tokens (min, max): 66.42 (2, 8.14M) Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/swedish-dynaword.imagetext-generation1B<n<10B3 likes1.4k downloads15d agoHugging Face04bkai-foundation-models /BKAINewsCorpus Dataset Card for "BKAINewsCorpus" The Binhvq News Corpus, a widely used dataset featuring approximately 20 million articles from diverse sources, received its last update in May 2021. To enhance this collection, we gathered an additional 10 million articles up until November 2023. By integrating these newly acquired articles with the existing Binhvq News Corpus, we have created an extensive Vietnamese News Corpus comprising about 32M articles. Subsequent fuzzy deduplication was… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/BKAINewsCorpus.text10M<n<100M14 likes1.1k downloads3y agoHugging Face05danish-foundation-models /dutch-dynaword 🧨 Dutch Dynaword Version 1.0.1 (Changelog) Language nld, Nederlands, Dutch License Openly Licensed, See the respective dataset Models For model trained used this data see danish-foundation-models Contact If you have question about this project please create an issue here Dataset Description Number of samples: 14.45M Number of tokens (Llama 3): 37.89B Average document length in tokens (min, max): 2.62K (2, 5.45M) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/dutch-dynaword.imagetext-generation10M<n<100M3 likes835 downloads13d agoHugging Face06danish-foundation-models /icelandic-dynaword 🧨 Icelandic Dynaword Version 0.0.15 (Changelog) Language Icelandic (is, isl) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 39.85M Number of tokens (Llama 3): 2.67B Average document length in tokens (min, max): 66.98 (3, 1.03M) Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dynaword.imagetext-generation100M<n<1B4 likes818 downloads16d agoHugging Face07danish-foundation-models /multilingual-gsm-symbolic Multilingual GSM-Symbolic Multilingual GSM-Symbolic is a benchmark for evaluating arithmetic reasoning in large language models across multiple languages. It extends Apple's GSM-Symbolic approach by providing symbolic templates that generate thousands of structurally equivalent but numerically distinct math problems. Templates and generation are handled by the multilingual-gsm-symbolic package. The dataset lets you test whether a model genuinely understands a problem or merely… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/multilingual-gsm-symbolic.text10K<n<100K3 likes730 downloads2mo agoHugging Face08danish-foundation-models /faroese-dynaword 🧨 Faroese Dynaword Version 0.0.7 (Changelog) Language Faroese (fo, fao) License Openly Licensed, See the respective dataset Models Currently there are no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 405.81K Number of tokens (Llama 3): 45.40M Average document length in tokens (min, max): 111.87 (2, 109.50K) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/faroese-dynaword.imagetext-generation1M<n<10M3 likes588 downloads7d agoHugging Face09bkai-foundation-models /NewsSapoVietnamese NewsSapo Dataset The Vietnamese NewsSapo dataset was constructed to train sentence/passage embeddings. Our dataset is structured in a "title-abstract-contents" format, where each news article is represented by a tuple of (title, abstract, content). The content is the main text body of the article and has been processed to remove images, videos, and other non-textual elements. The dataset contains 31,728,183 triples. To build this dataset, we followed a two-step process: Step 1:… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/NewsSapo.textsummarization1M<n<10M6 likes563 downloads3y agoHugging Face10bkai-foundation-models /vi-alpaca 🇻🇳 Vietnamese Alpaca Dataset This dataset is especially designed for Vietnamese based on the idea from Stanford Alpaca and Self-Instruct paper. The motivation behind the creation of this dataset stems from the hope to contribute high-quality dataset to Vietnamese commnunity to train language models. To construct this dataset, we follow a two-step process: Step 1: Manually create Vietnamese seed tasks We employ the methodology outlined in the Self-Instruct paper we meticulously… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/vi-alpaca.text10K<n<100K25 likes510 downloads3y agoHugging Face11danish-foundation-models /danish-gigaword Danish Gigaword Corpus Version: 1.0.0 License: See the respective dataset Dataset Summary The Danish Gigaword Corpus contains text spanning several domains and forms. This version does not include the sections containing tweets ("General Discussions" and "Parliament Elections"), "danavis", "Common Crawl" and "OpenSubtitles" due to potential privacy, quality and copyright concerns. Loading the dataset from datasets import load_dataset name =… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-gigaword.texttext-generation100K<n<1M9 likes468 downloads2y agoHugging Face12danish-foundation-models /multi-ifeval MultiIFEval This dataset is an instruction-following dataset for 300+ languages, translated and localised from the English IFEval dataset. Dataset Details Dataset Description All samples come from the English IFEval dataset, and we translate and localise with Gemini-3-flash-preview. When translating and localising samples, we also include a random Wikipedia article in the target language, both to give some context for localisation, but also to… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/multi-ifeval.text100K<n<1M2 likes396 downloads2mo agoHugging Face13genbio-ai /foundation-models-perturbationData for the paper "Foundation Models Improve Perturbation Response Prediction" as described on GitHub. text100K<n<1M0 likes335 downloads7mo agoHugging Face14foundation-models /milp-instances-parquet MILP instances (Parquet) Competition-style instances packed as Zstd-compressed Parquet shards for partial downloads. Schema Column Type Description instance_id string Stem name (e.g. load_balancing_0) task string item_placement, load_balancing, or anonymous split string train or valid json_text string Raw contents of the sidecar .json mps_gz binary Bytes of the .mps.gz file Tasks are independent (separate folders / configs). Shards are named… See the full description on the dataset page: https://huggingface.co/datasets/foundation-models/milp-instances-parquet.text10K<n<100K0 likes287 downloads6mo agoHugging Face15danish-foundation-models /dala_gen_v3text1K<n<10K0 likes254 downloads5mo agoHugging Face16danish-foundation-models /norwegian-dyna-instruct 🧨 Norwegian dyna-instruct Version 0.1.0 (changelog) Languages Norwegian Bokmål (nob), Norwegian Nynorsk (nno), and English (eng) translation input License Mixed open licenses; see the table below Sources Five datasets (source cards) Dataset Description Number of samples: 14.40K Number of tokens (Llama 3): 6.27M Average conversation length in tokens (min, max): 435.63 (4, 8.92K) Average number of turns (min, max): 2.13 (2, 3)… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dyna-instruct.imagequestion-answering10K<n<100K0 likes233 downloads17d agoHugging Face17danish-foundation-models /ifeval-da IFEval-da This dataset is a translation of the English IFEval dataset, which was published in this paper and contains 541 prompts, each with a combination of one or more of 25 different constraints. The dataset was professionally translated and localised by expert native speakers. Dataset Details Translated by: Rasmus Larsen (rasmus.larsen@alexandra.dk), Nathalie Hau Sørensen (naha@hum.ku.dk) and Kenneth Enevoldsen (kenneth.enevoldsen@cas.au.dk) Funded by: Danish… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/ifeval-da.texttext-generationn<1K1 likes228 downloads7mo agoHugging Face18danish-foundation-models /faroese-dyna-instruct 🧨 Faroese dyna-instruct Version 0.1.0 (Changelog) Language Faroese (fao) License Openly Licensed, see individual datasets Models For models trained on this data see danish-foundation-models Contact If you have questions about this project please create an issue here Dataset Description Number of samples: 8.61K Number of tokens (Llama 3): 2.64M Average conversation length in tokens (min, max): 306.67 (98, 1.24K) Average number of… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/faroese-dyna-instruct.texttext-generation10K<n<100K1 likes222 downloads22d agoHugging Face19foundation-multimodal-models /DetailCaps-4870 DetailCaps-4870 Benchmark The detail image caption evaluation benchmark proposed in our paper Benchmarking and Improving Detail Image Caption. 🏠 Homepage | 📑 Paper | 🤗 Huggingface Datasets Overview We curate 4870 images from various datasets, accompanying with ground truth detail captions generated by GPT-4V, Gemini-1.5-Pro and GPT-4O for evaluation. We also provide captions generated by three open-source LVLMs, which are LLaVA-1.5, CogVLM and ShareCaptioner, as well… See the full description on the dataset page: https://huggingface.co/datasets/foundation-multimodal-models/DetailCaps-4870.text1K<n<10K15 likes216 downloads2y agoHugging Face20danish-foundation-models /icelandic-dyna-instruct 🧨 Icelandic dyna-instruct Version 0.1.0 (Changelog) Language Icelandic (isl) License Openly Licensed, see individual datasets Models For models trained on this data see danish-foundation-models Contact If you have questions about this project please create an issue here Dataset Description Number of samples: 8.11K Number of tokens (Llama 3): 7.09M Average conversation length in tokens (min, max): 874.89 (182, 1.39K) Average number… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dyna-instruct.texttext-generation10K<n<100K1 likes159 downloads22d agoHugging Face21foundation-multimodal-models /ConBench_Dimage1K<n<10K0 likes97 downloads2y agoHugging Face22danish-foundation-models /nasjonalt-vitenarkiv Nasjonalt vitenarkiv Open-access documents from NVA (Nasjonalt vitenarkiv), the joint national repository where Norwegian research institutions publish their output: master's and PhD theses, journal articles, and technical and research reports. Subjects span the disciplines - marine science, forestry, archaeology, education, public health, engineering - and most documents are recent. Each row is one PDF: the original file exactly as published, the text extracted from it, and the… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/nasjonalt-vitenarkiv.documenttext-generationn<1K1 likes82 downloads2mo agoHugging Face23danish-foundation-models /dala DaLA: Danish Linguistic Acceptability Evaluation Dataset DaLA (paper) is a benchmark dataset for linguistic acceptability judgment in Danish, designed to evaluate how well NLP models, especially large language models (LLMs), understand grammaticality in real-world Danish sentences. The dataset extends previous resources by introducing a broader and more realistic set of error types and providing data splits suitable for evaluation via few-shot or finetuning. 🔗… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/dala.texttext-classification1K<n<10K0 likes75 downloads7mo agoHugging Face24foundation-models /sld-imagesimage10K<n<100K0 likes64 downloads8mo agoHugging Face25danish-foundation-models /dala_large DaLA: Danish Linguistic Acceptability Evaluation Dataset DaLA (paper) is a benchmark dataset for linguistic acceptability judgment in Danish, designed to evaluate how well NLP models, especially large language models (LLMs), understand grammaticality in real-world Danish sentences. The dataset extends previous resources by introducing a broader and more realistic set of error types and providing data splits suitable for evaluation via few-shot or finetuning. 🔗… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/dala_large.texttext-classification1K<n<10K0 likes48 downloads7mo agoHugging Face26danish-foundation-models /dala_medium DaLA: Danish Linguistic Acceptability Evaluation Dataset DaLA (paper) is a benchmark dataset for linguistic acceptability judgment in Danish, designed to evaluate how well NLP models, especially large language models (LLMs), understand grammaticality in real-world Danish sentences. The dataset extends previous resources by introducing a broader and more realistic set of error types and providing data splits suitable for evaluation via few-shot or finetuning. 🔗… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/dala_medium.texttext-classification1K<n<10K0 likes46 downloads7mo agoHugging Face27bkai-foundation-models /vi-self-chat-sharegpt-format 🇻🇳 Vietnamese Self-Chat Dataset This dataset is designed to enhance the model's ability to engage in multi-turn conversations with humans. To construct this dataset, we follow a two-step process: Step 1: Instruction Generation We employ the methodology outlined in the Self-Instruct paper to craft a diverse set of instructions. This paper serves as a guide for aligning pretrained language models with specific instructions, providing a structured foundation for subsequent dialogue… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/vi-self-chat-sharegpt-format.text10K<n<100K13 likes42 downloads3y agoHugging Face28bkai-foundation-models /NewsCategorygated Overview The dataset is collected from the Vnexpress news website and is extracted for clustering tasks. (596524 samples) Data Format The data consists of 5 fields: id: The index of the article. title: The title of the article. sapo: The summary of the article. content: The main content of the article. label: The topic of the article. Article Topics The articles are categorized into 21 topics, including: 'Ngôi Sao' (Celebrities) 'Thế giới' (World) 'Giải trí giới trẻ'… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/NewsCategory.text100K<n<1M4 likes33 downloads3y agoHugging Face29danish-foundation-models /laerebogengated Lærebogen An instruction-following dataset for Danish. This dataset features 5 million examples of multi-turn conversations in Danish, designed to train instruction-following models, with a commercially usable license. Dataset Structure All examples in the dataset are structured as follows: { "messages": [ { "role": "user", "content": "(...)" }, { "role": "assistant", "content": "(...)" }, { "role": "user", "content": "(...)" }, (...) { "role":… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/laerebogen.texttext-generation1M<n<10M1 likes29 downloads6mo agoHugging Face30bkai-foundation-models /vi-alpaca-input-output-format 🇻🇳 Vietnamese modified Alpaca Dataset This dataset is especially designed for Vietnamese based on the idea from Stanford Alpaca, Self-Instruct paper and Chinese LLaMA. The motivation behind the creation of this dataset stems from the hope to contribute high-quality dataset to Vietnamese commnunity to train language models. To construct this dataset, we follow a two-step process: Step 1: Manually create Vietnamese seed tasks We employ the methodology outlined in the Self-Instruct… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/vi-alpaca-input-output-format.text10K<n<100K7 likes28 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.