datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
danish-dynaword
🧨 Danish Dynaword
Version
1.2.23 (Changelog)
Language
dan, dansk, Danish
License
Openly Licensed, See the respective dataset
Models
For model trained used this data see danish-foundation-models
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 7.40M
Number of tokens (Llama 3): 9.81B
Average document length in tokens (min, max): 1.33K (2, 19.46M)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-dynaword.norwegian-dynaword
🧨 Norwegian Dynaword
Version
0.0.18 (Changelog)
Language
Norwegian (no, nor), including Bokmål (nb, nob) and Nynorsk (nn, nno)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 4.47M
Number of tokens (Llama 3): 9.98B
Average document length in tokens (min… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dynaword.swedish-dynaword
🧨 Swedish Dynaword
Version
0.0.13 (Changelog)
Language
Swedish (sv, swe)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 547.06M
Number of tokens (Llama 3): 36.34B
Average document length in tokens (min, max): 66.42 (2, 8.14M)
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/swedish-dynaword.BKAINewsCorpus
Dataset Card for "BKAINewsCorpus"
The Binhvq News Corpus, a widely used dataset featuring approximately 20 million articles from diverse sources, received its last update in May 2021. To enhance this collection, we gathered an additional 10 million articles up until November 2023. By integrating these newly acquired articles with the existing Binhvq News Corpus, we have created an extensive Vietnamese News Corpus comprising about 32M articles. Subsequent fuzzy deduplication was… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/BKAINewsCorpus.dutch-dynaword
🧨 Dutch Dynaword
Version
1.0.1 (Changelog)
Language
nld, Nederlands, Dutch
License
Openly Licensed, See the respective dataset
Models
For model trained used this data see danish-foundation-models
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 14.45M
Number of tokens (Llama 3): 37.89B
Average document length in tokens (min, max): 2.62K (2, 5.45M)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/dutch-dynaword.icelandic-dynaword
🧨 Icelandic Dynaword
Version
0.0.15 (Changelog)
Language
Icelandic (is, isl)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 39.85M
Number of tokens (Llama 3): 2.67B
Average document length in tokens (min, max): 66.98 (3, 1.03M)
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dynaword.multilingual-gsm-symbolic
Multilingual GSM-Symbolic
Multilingual GSM-Symbolic is a benchmark for evaluating arithmetic reasoning in large language models across multiple languages. It extends Apple's GSM-Symbolic approach by providing symbolic templates that generate thousands of structurally equivalent but numerically distinct math problems. Templates and generation are handled by the multilingual-gsm-symbolic package.
The dataset lets you test whether a model genuinely understands a problem or merely… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/multilingual-gsm-symbolic.foundation-model-results-simulation-consortiafoundation-model-data-experimentalNewsSapoVietnamese NewsSapo Dataset
The Vietnamese NewsSapo dataset was constructed to train sentence/passage embeddings. Our dataset is structured in a "title-abstract-contents" format, where each news article is represented by a tuple of (title, abstract, content). The content is the main text body of the article and has been processed to remove images, videos, and other non-textual elements. The dataset contains 31,728,183 triples.
To build this dataset, we followed a two-step process:
Step 1:… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/NewsSapo.faroese-dynaword
🧨 Faroese Dynaword
Version
0.0.7 (Changelog)
Language
Faroese (fo, fao)
License
Openly Licensed, See the respective dataset
Models
Currently there are no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 405.81K
Number of tokens (Llama 3): 45.40M
Average document length in tokens (min, max): 111.87 (2, 109.50K)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/faroese-dynaword.vi-alpaca
🇻🇳 Vietnamese Alpaca Dataset
This dataset is especially designed for Vietnamese based on the idea from Stanford Alpaca and Self-Instruct paper. The motivation behind the creation of this dataset stems from the hope to contribute high-quality dataset to Vietnamese commnunity to train language models.
To construct this dataset, we follow a two-step process:
Step 1: Manually create Vietnamese seed tasks
We employ the methodology outlined in the Self-Instruct paper we meticulously… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/vi-alpaca.danish-gigaword
Danish Gigaword Corpus
Version: 1.0.0
License: See the respective dataset
Dataset Summary
The Danish Gigaword Corpus contains text spanning several domains and forms. This version does not include the sections containing tweets ("General Discussions" and "Parliament Elections"), "danavis", "Common Crawl" and "OpenSubtitles" due to potential privacy, quality and copyright concerns.
Loading the dataset
from datasets import load_dataset
name =… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-gigaword.multi-ifeval
MultiIFEval
This dataset is an instruction-following dataset for 300+ languages, translated and localised from the English IFEval dataset.
Dataset Details
Dataset Description
All samples come from the English IFEval dataset, and we translate and localise with Gemini-3-flash-preview.
When translating and localising samples, we also include a random Wikipedia article in the target language, both to give some context for localisation, but also to… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/multi-ifeval.foundation-models-perturbationData for the paper "Foundation Models Improve Perturbation Response Prediction" as described on GitHub.
dala_gen_v3milp-instances-parquet
MILP instances (Parquet)
Competition-style instances packed as Zstd-compressed Parquet shards for partial downloads.
Schema
Column
Type
Description
instance_id
string
Stem name (e.g. load_balancing_0)
task
string
item_placement, load_balancing, or anonymous
split
string
train or valid
json_text
string
Raw contents of the sidecar .json
mps_gz
binary
Bytes of the .mps.gz file
Tasks are independent (separate folders / configs). Shards are named… See the full description on the dataset page: https://huggingface.co/datasets/foundation-models/milp-instances-parquet.ifeval-da
IFEval-da
This dataset is a translation of the English IFEval dataset,
which was published in this paper and contains 541 prompts,
each with a combination of one or more of 25 different constraints. The dataset was professionally
translated and localised by expert native speakers.
Dataset Details
Translated by: Rasmus Larsen (rasmus.larsen@alexandra.dk), Nathalie Hau Sørensen (naha@hum.ku.dk) and Kenneth Enevoldsen (kenneth.enevoldsen@cas.au.dk)
Funded by: Danish… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/ifeval-da.norwegian-dyna-instruct
🧨 Norwegian dyna-instruct
Version
0.1.0 (changelog)
Languages
Norwegian Bokmål (nob), Norwegian Nynorsk (nno), and English (eng) translation input
License
Mixed open licenses; see the table below
Sources
Five datasets (source cards)
Dataset Description
Number of samples: 14.40K
Number of tokens (Llama 3): 6.27M
Average conversation length in tokens (min, max): 435.63 (4, 8.92K)
Average number of turns (min, max): 2.13 (2, 3)… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dyna-instruct.faroese-dyna-instruct
🧨 Faroese dyna-instruct
Version
0.1.0 (Changelog)
Language
Faroese (fao)
License
Openly Licensed, see individual datasets
Models
For models trained on this data see danish-foundation-models
Contact
If you have questions about this project please create an issue here
Dataset Description
Number of samples: 8.61K
Number of tokens (Llama 3): 2.64M
Average conversation length in tokens (min, max): 306.67 (98, 1.24K)
Average number of… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/faroese-dyna-instruct.DetailCaps-4870
DetailCaps-4870 Benchmark
The detail image caption evaluation benchmark proposed in our paper Benchmarking and Improving Detail Image Caption.
🏠 Homepage | 📑 Paper | 🤗 Huggingface Datasets
Overview
We curate 4870 images from various datasets, accompanying with ground truth detail captions generated by GPT-4V, Gemini-1.5-Pro and GPT-4O for evaluation.
We also provide captions generated by three open-source LVLMs, which are LLaVA-1.5, CogVLM and ShareCaptioner, as well… See the full description on the dataset page: https://huggingface.co/datasets/foundation-multimodal-models/DetailCaps-4870.foundation-model-data-simulationicelandic-dyna-instruct
🧨 Icelandic dyna-instruct
Version
0.1.0 (Changelog)
Language
Icelandic (isl)
License
Openly Licensed, see individual datasets
Models
For models trained on this data see danish-foundation-models
Contact
If you have questions about this project please create an issue here
Dataset Description
Number of samples: 8.11K
Number of tokens (Llama 3): 7.09M
Average conversation length in tokens (min, max): 874.89 (182, 1.39K)
Average number… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dyna-instruct.crosslingual
VNLAWQC, VNSynLawQC: A Vietnamese Legal Retrieval Dataset
VNLAWQC, is sourced from the Vietnamese Law Library (VLL). The VLL contains articles that address questions spanning multiple aspects of the legal domain. Each article provides an answer supported by one or more legal documents, with hyperlinks directing to the corresponding documents.
VNSynLawQC is augmented based on law documents in VNLAWQC using Llama-3-70B.
Dataset Composition
The dataset consists of query… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/crosslingual.ConBench_Dgolden-batch-sentinel-data
Golden Batch Sentinel Data
Benchmark datasets for process monitoring and fault detection in batch manufacturing.
Datasets
IndPenSim (Industrial Penicillin Simulation)
A 100,000L fermentation simulation with 100 batches and rich multivariate signals.
Source: Mendeley Data
Paper: Modern day monitoring and control challenges...
Batches: 100 (90 normal, 10 faulty)
Variables: 37 process variables (Raman spectra excluded for efficiency)
Time resolution: 0.2 hours… See the full description on the dataset page: https://huggingface.co/datasets/foundation-models/golden-batch-sentinel-data.social-media-campaigns
video-factory
A reusable pipeline for producing short, vertical (1080×1350, 4:5) explainer
videos that pair narrated avatar clips with self-contained HTML/CSS
kinetic-typography animations. Built for a daily publishing cadence to
LinkedIn / YouTube.
The HTML is the source of truth — it is meant to be hand-edited. Everything
else (per-section MP4s, the concatenated final.mp4) is regenerated from it.
Layout
video-factory/
├── engine/ # shared… See the full description on the dataset page: https://huggingface.co/datasets/foundation-models/social-media-campaigns.synthetic-values-model-charter
value_units.jsonl is the individual parsed values from the model charter.
scenarios.jsonl is invididual hypothetical scenarios based on the values in values_units.jsonl
sft_*.jsonl generated accepted responses.
dpo_*.jsonl generated accepted+rejected responses.
nasjonalt-vitenarkiv
Nasjonalt vitenarkiv
Open-access documents from NVA (Nasjonalt vitenarkiv), the joint national
repository where Norwegian research institutions publish their output: master's and PhD theses,
journal articles, and technical and research reports. Subjects span the disciplines - marine
science, forestry, archaeology, education, public health, engineering - and most documents are
recent.
Each row is one PDF: the original file exactly as published, the text extracted from it, and the… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/nasjonalt-vitenarkiv.TCGA_foundation_model_features
