CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01danish-foundation-models /danish-dynaword 🧨 Danish Dynaword Version 1.2.23 (Changelog) Language dan, dansk, Danish License Openly Licensed, See the respective dataset Models For model trained used this data see danish-foundation-models Contact If you have question about this project please create an issue here Dataset Description Number of samples: 7.40M Number of tokens (Llama 3): 9.81B Average document length in tokens (min, max): 1.33K (2, 19.46M) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-dynaword.imagetext-generation10M<n<100M22 likes11k downloads21d agoHugging Face02danish-foundation-models /norwegian-dynaword 🧨 Norwegian Dynaword Version 0.0.18 (Changelog) Language Norwegian (no, nor), including Bokmål (nb, nob) and Nynorsk (nn, nno) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 4.47M Number of tokens (Llama 3): 9.98B Average document length in tokens (min… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dynaword.imagetext-generation10M<n<100M7 likes1.9k downloads15d agoHugging Face03danish-foundation-models /swedish-dynaword 🧨 Swedish Dynaword Version 0.0.13 (Changelog) Language Swedish (sv, swe) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 547.06M Number of tokens (Llama 3): 36.34B Average document length in tokens (min, max): 66.42 (2, 8.14M) Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/swedish-dynaword.imagetext-generation1B<n<10B3 likes1.4k downloads13d agoHugging Face04bkai-foundation-models /BKAINewsCorpus Dataset Card for "BKAINewsCorpus" The Binhvq News Corpus, a widely used dataset featuring approximately 20 million articles from diverse sources, received its last update in May 2021. To enhance this collection, we gathered an additional 10 million articles up until November 2023. By integrating these newly acquired articles with the existing Binhvq News Corpus, we have created an extensive Vietnamese News Corpus comprising about 32M articles. Subsequent fuzzy deduplication was… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/BKAINewsCorpus.text10M<n<100M14 likes1.1k downloads3y agoHugging Face05danish-foundation-models /dutch-dynaword 🧨 Dutch Dynaword Version 1.0.1 (Changelog) Language nld, Nederlands, Dutch License Openly Licensed, See the respective dataset Models For model trained used this data see danish-foundation-models Contact If you have question about this project please create an issue here Dataset Description Number of samples: 14.45M Number of tokens (Llama 3): 37.89B Average document length in tokens (min, max): 2.62K (2, 5.45M) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/dutch-dynaword.imagetext-generation10M<n<100M3 likes824 downloads12d agoHugging Face06danish-foundation-models /icelandic-dynaword 🧨 Icelandic Dynaword Version 0.0.15 (Changelog) Language Icelandic (is, isl) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 39.85M Number of tokens (Llama 3): 2.67B Average document length in tokens (min, max): 66.98 (3, 1.03M) Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dynaword.imagetext-generation100M<n<1B4 likes806 downloads15d agoHugging Face07danish-foundation-models /multilingual-gsm-symbolic Multilingual GSM-Symbolic Multilingual GSM-Symbolic is a benchmark for evaluating arithmetic reasoning in large language models across multiple languages. It extends Apple's GSM-Symbolic approach by providing symbolic templates that generate thousands of structurally equivalent but numerically distinct math problems. Templates and generation are handled by the multilingual-gsm-symbolic package. The dataset lets you test whether a model genuinely understands a problem or merely… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/multilingual-gsm-symbolic.text10K<n<100K3 likes717 downloads2mo agoHugging Face08bkai-foundation-models /NewsSapoVietnamese NewsSapo Dataset The Vietnamese NewsSapo dataset was constructed to train sentence/passage embeddings. Our dataset is structured in a "title-abstract-contents" format, where each news article is represented by a tuple of (title, abstract, content). The content is the main text body of the article and has been processed to remove images, videos, and other non-textual elements. The dataset contains 31,728,183 triples. To build this dataset, we followed a two-step process: Step 1:… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/NewsSapo.textsummarization1M<n<10M6 likes571 downloads3y agoHugging Face09danish-foundation-models /faroese-dynaword 🧨 Faroese Dynaword Version 0.0.7 (Changelog) Language Faroese (fo, fao) License Openly Licensed, See the respective dataset Models Currently there are no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 405.81K Number of tokens (Llama 3): 45.40M Average document length in tokens (min, max): 111.87 (2, 109.50K) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/faroese-dynaword.imagetext-generation1M<n<10M3 likes516 downloads6d agoHugging Face10bkai-foundation-models /vi-alpaca 🇻🇳 Vietnamese Alpaca Dataset This dataset is especially designed for Vietnamese based on the idea from Stanford Alpaca and Self-Instruct paper. The motivation behind the creation of this dataset stems from the hope to contribute high-quality dataset to Vietnamese commnunity to train language models. To construct this dataset, we follow a two-step process: Step 1: Manually create Vietnamese seed tasks We employ the methodology outlined in the Self-Instruct paper we meticulously… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/vi-alpaca.text10K<n<100K25 likes513 downloads3y agoHugging Face11danish-foundation-models /danish-gigaword Danish Gigaword Corpus Version: 1.0.0 License: See the respective dataset Dataset Summary The Danish Gigaword Corpus contains text spanning several domains and forms. This version does not include the sections containing tweets ("General Discussions" and "Parliament Elections"), "danavis", "Common Crawl" and "OpenSubtitles" due to potential privacy, quality and copyright concerns. Loading the dataset from datasets import load_dataset name =… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-gigaword.texttext-generation100K<n<1M9 likes468 downloads2y agoHugging Face12danish-foundation-models /multi-ifeval MultiIFEval This dataset is an instruction-following dataset for 300+ languages, translated and localised from the English IFEval dataset. Dataset Details Dataset Description All samples come from the English IFEval dataset, and we translate and localise with Gemini-3-flash-preview. When translating and localising samples, we also include a random Wikipedia article in the target language, both to give some context for localisation, but also to… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/multi-ifeval.text100K<n<1M2 likes361 downloads2mo agoHugging Face13genbio-ai /foundation-models-perturbationData for the paper "Foundation Models Improve Perturbation Response Prediction" as described on GitHub. text100K<n<1M0 likes335 downloads7mo agoHugging Face14danish-foundation-models /dala_gen_v3text1K<n<10K0 likes291 downloads5mo agoHugging Face15foundation-models /milp-instances-parquet MILP instances (Parquet) Competition-style instances packed as Zstd-compressed Parquet shards for partial downloads. Schema Column Type Description instance_id string Stem name (e.g. load_balancing_0) task string item_placement, load_balancing, or anonymous split string train or valid json_text string Raw contents of the sidecar .json mps_gz binary Bytes of the .mps.gz file Tasks are independent (separate folders / configs). Shards are named… See the full description on the dataset page: https://huggingface.co/datasets/foundation-models/milp-instances-parquet.text10K<n<100K0 likes286 downloads5mo agoHugging Face16danish-foundation-models /ifeval-da IFEval-da This dataset is a translation of the English IFEval dataset, which was published in this paper and contains 541 prompts, each with a combination of one or more of 25 different constraints. The dataset was professionally translated and localised by expert native speakers. Dataset Details Translated by: Rasmus Larsen (rasmus.larsen@alexandra.dk), Nathalie Hau Sørensen (naha@hum.ku.dk) and Kenneth Enevoldsen (kenneth.enevoldsen@cas.au.dk) Funded by: Danish… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/ifeval-da.texttext-generationn<1K1 likes231 downloads7mo agoHugging Face17danish-foundation-models /norwegian-dyna-instruct 🧨 Norwegian dyna-instruct Version 0.1.0 (changelog) Languages Norwegian Bokmål (nob), Norwegian Nynorsk (nno), and English (eng) translation input License Mixed open licenses; see the table below Sources Five datasets (source cards) Dataset Description Number of samples: 14.40K Number of tokens (Llama 3): 6.27M Average conversation length in tokens (min, max): 435.63 (4, 8.92K) Average number of turns (min, max): 2.13 (2, 3)… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dyna-instruct.imagequestion-answering10K<n<100K0 likes230 downloads15d agoHugging Face18danish-foundation-models /faroese-dyna-instruct 🧨 Faroese dyna-instruct Version 0.1.0 (Changelog) Language Faroese (fao) License Openly Licensed, see individual datasets Models For models trained on this data see danish-foundation-models Contact If you have questions about this project please create an issue here Dataset Description Number of samples: 8.61K Number of tokens (Llama 3): 2.64M Average conversation length in tokens (min, max): 306.67 (98, 1.24K) Average number of… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/faroese-dyna-instruct.texttext-generation10K<n<100K1 likes217 downloads20d agoHugging Face19foundation-multimodal-models /DetailCaps-4870 DetailCaps-4870 Benchmark The detail image caption evaluation benchmark proposed in our paper Benchmarking and Improving Detail Image Caption. 🏠 Homepage | 📑 Paper | 🤗 Huggingface Datasets Overview We curate 4870 images from various datasets, accompanying with ground truth detail captions generated by GPT-4V, Gemini-1.5-Pro and GPT-4O for evaluation. We also provide captions generated by three open-source LVLMs, which are LLaVA-1.5, CogVLM and ShareCaptioner, as well… See the full description on the dataset page: https://huggingface.co/datasets/foundation-multimodal-models/DetailCaps-4870.text1K<n<10K15 likes193 downloads2y agoHugging Face20danish-foundation-models /icelandic-dyna-instruct 🧨 Icelandic dyna-instruct Version 0.1.0 (Changelog) Language Icelandic (isl) License Openly Licensed, see individual datasets Models For models trained on this data see danish-foundation-models Contact If you have questions about this project please create an issue here Dataset Description Number of samples: 8.11K Number of tokens (Llama 3): 7.09M Average conversation length in tokens (min, max): 874.89 (182, 1.39K) Average number… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dyna-instruct.texttext-generation10K<n<100K1 likes158 downloads20d agoHugging Face21bkai-foundation-models /crosslingual VNLAWQC, VNSynLawQC: A Vietnamese Legal Retrieval Dataset VNLAWQC, is sourced from the Vietnamese Law Library (VLL). The VLL contains articles that address questions spanning multiple aspects of the legal domain. Each article provides an answer supported by one or more legal documents, with hyperlinks directing to the corresponding documents. VNSynLawQC is augmented based on law documents in VNLAWQC using Llama-3-70B. Dataset Composition The dataset consists of query… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/crosslingual.feature-extraction3 likes143 downloads2y agoHugging Face22foundation-multimodal-models /ConBench_Dimage1K<n<10K0 likes97 downloads2y agoHugging Face23foundation-models /golden-batch-sentinel-data Golden Batch Sentinel Data Benchmark datasets for process monitoring and fault detection in batch manufacturing. Datasets IndPenSim (Industrial Penicillin Simulation) A 100,000L fermentation simulation with 100 batches and rich multivariate signals. Source: Mendeley Data Paper: Modern day monitoring and control challenges... Batches: 100 (90 normal, 10 faulty) Variables: 37 process variables (Raman spectra excluded for efficiency) Time resolution: 0.2 hours… See the full description on the dataset page: https://huggingface.co/datasets/foundation-models/golden-batch-sentinel-data.tabulartime-series-forecasting10M<n<100M0 likes97 downloads8mo agoHugging Face24foundation-models /social-media-campaigns video-factory A reusable pipeline for producing short, vertical (1080×1350, 4:5) explainer videos that pair narrated avatar clips with self-contained HTML/CSS kinetic-typography animations. Built for a daily publishing cadence to LinkedIn / YouTube. The HTML is the source of truth — it is meant to be hand-edited. Everything else (per-section MP4s, the concatenated final.mp4) is regenerated from it. Layout video-factory/ ├── engine/ # shared… See the full description on the dataset page: https://huggingface.co/datasets/foundation-models/social-media-campaigns.0 likes89 downloads28d agoHugging Face25danish-foundation-models /synthetic-values-model-charter value_units.jsonl is the individual parsed values from the model charter. scenarios.jsonl is invididual hypothetical scenarios based on the values in values_units.jsonl sft_*.jsonl generated accepted responses. dpo_*.jsonl generated accepted+rejected responses. 0 likes85 downloads27d agoHugging Face26danish-foundation-models /nasjonalt-vitenarkiv Nasjonalt vitenarkiv Open-access documents from NVA (Nasjonalt vitenarkiv), the joint national repository where Norwegian research institutions publish their output: master's and PhD theses, journal articles, and technical and research reports. Subjects span the disciplines - marine science, forestry, archaeology, education, public health, engineering - and most documents are recent. Each row is one PDF: the original file exactly as published, the text extracted from it, and the… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/nasjonalt-vitenarkiv.documenttext-generationn<1K1 likes83 downloads2mo agoHugging Face27danish-foundation-models /dala DaLA: Danish Linguistic Acceptability Evaluation Dataset DaLA (paper) is a benchmark dataset for linguistic acceptability judgment in Danish, designed to evaluate how well NLP models, especially large language models (LLMs), understand grammaticality in real-world Danish sentences. The dataset extends previous resources by introducing a broader and more realistic set of error types and providing data splits suitable for evaluation via few-shot or finetuning. 🔗… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/dala.texttext-classification1K<n<10K0 likes75 downloads7mo agoHugging Face28foundation-models /sld-imagesimage10K<n<100K0 likes64 downloads8mo agoHugging Face29foundation-models /evidentia-public0 likes56 downloads3mo agoHugging Face30foundation-models /imagesimagen<1K0 likes55 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.